Welcome to the VBNN Digital Library
Unlock a Vast Knowledge Ecosystem
Featuring over 30,000 books, academic papers, illustrations, and expert insights—continuously updated to support your research and professional growth.
Welcome to our library!
Here, you will find an exclusive collection created 100% by our own faculty, meaning you will not find these resources anywhere else. Over the last 20 years, our team has written much more than what is currently online, and we are actively working to upload our complete back catalog. We update our platform regularly, so be sure to check back from time to time. If you ever need help finding a specific resource, you can always contact us!
Maximize Your Access
Log in to instantly view and download tailored resources directly aligned with your specific program and curriculum.
Ready to begin? Sign in above to explore your personalized dashboard.
Please note: Login is only possible using your institutional email address; otherwise, the system will not recognize your account.
VBNN Library AI
Introducing our fully integrated Library AI. Designed to support your research, you may submit inquiries in any language and receive precise, evidence-based responses drawn exclusively from our published scholarly articles and textbooks.
Search...
Latest Publications:
Search this site
Results found for empty search
- Laboratory Information Management Systems | LIMS (Deployment and Compliance)
Download the Book (PDF): Introduction Every regulated laboratory eventually receives the same question in one form or another. An inspector, an auditor, a sponsor, or a sceptical colleague points to a number in a report and asks how the laboratory knows it is true. Not whether the number is plausible, and not whether the scientist who produced it is trustworthy, but how the laboratory can show, from its own records, which sample the number came from, who handled that sample, which instrument measured it, what settings were used, whether anything was changed afterwards, and why. The laboratory either reconstructs that history from evidence or it does not. Everything else about compliance follows from that moment. A Laboratory Information Management System, or LIMS, is the tool most laboratories now use to make that reconstruction possible. At its simplest a LIMS is a database with a workflow engine attached: it registers samples, assigns them identifiers, schedules the tests to be performed on them, records results against specifications, and produces reports. In practice it becomes the spine of the laboratory's record-keeping, the place where sample identity, custody, analytical results, instrument links and approvals meet. That is why regulators take such an interest in it, and why a poorly configured LIMS is so much more dangerous than a poorly kept notebook. A notebook fails one page at a time. A LIMS fails systematically, and it does so with the authority of a computer printout. This booklet is about deploying and keeping a LIMS in the particular circumstances of academic research laboratories that work under Good Manufacturing Practice (GMP) or Good Laboratory Practice (GLP). Those circumstances are more common than they once were. Universities and university hospitals now run manufacturing facilities for cell and gene therapies, radiopharmaceuticals and early-phase investigational medicines. Academic toxicology, pharmacology and analytical cores conduct GLP studies for regulatory submissions. Research institutes operate biobanks and contract testing services that are held to recognised quality standards. In each case a laboratory that grew up with the habits of discovery science, in which the scientist is the authority and the notebook is personal property, has to adopt the habits of regulated science, in which the record is the authority and belongs to the organisation. The transition is not mainly about software. It is about accepting a different theory of evidence. In discovery research a result is credible because it can be repeated. In regulated work a result is credible because its history can be checked. A LIMS is valuable only to the extent that it makes that history checkable, and it can just as easily obscure it. A system that permits shared logins, lets analysts overwrite results without a trace, holds instrument data in a spreadsheet that someone retyped, or has not been tested since the last upgrade does not make a laboratory compliant. It makes its non-compliance harder to see. The argument of this booklet The argument running through the chapters is simple to state and demanding to practise. A LIMS does not confer compliance. Compliance is a property of the evidence a laboratory can produce, and a LIMS earns its place only when it is configured, validated and maintained so that an independent person could reconstruct what happened to any sample and any result from the records alone. Configuration, validation and maintenance are therefore not three separate projects handled by three separate groups. They are one continuous discipline: designing the system around the reconstruction it must support, proving that it supports it, and keeping it that way while people, instruments, software versions and regulations all change underneath it. Holding to that single test has practical consequences that the chapters work through. It explains why the choice of sample identifier matters more than the choice of vendor. It explains why automated instrument capture is worth its considerable cost, and why a poorly designed interface can be worse than none. It explains why chain of custody is a set of recorded events rather than a signature on a form, and why validation effort should be concentrated on the functions that carry evidential weight rather than spread evenly over every screen. It explains why audit trail review, periodic review and change control are where most academic deployments quietly fail, long after the go-live celebration. Who this is for and what it covers The intended reader is a scientist, quality professional, facility manager or research computing specialist who has been handed responsibility for a LIMS in a regulated academic setting, or is about to be. No prior knowledge of the regulations is assumed, but the reader is expected to know what a laboratory does and to be comfortable with the idea of a database. The booklet does not recommend particular products. Commercial and open-source systems differ in cost, flexibility and support, but the principles that decide whether a deployment succeeds are the same across them, and product features change faster than books are printed. Chapter 1 sets out what a LIMS is for in evidential terms and why academic laboratories find the shift to regulated record-keeping difficult. Chapter 2 maps the regulatory ground: the United States Food and Drug Administration's rules on electronic records and on GLP and GMP, the European Union's Annex 11, the Organisation for Economic Co-operation and Development's GLP principles and advisory documents, and the data integrity guidance that has grown up around them since 2015. Chapter 3 treats configuration, the design of the sample model, workflows, specifications and user roles, including the problem of segregating duties in a group of six people. Chapter 4 is about identity: barcodes, labels and identifier schemes, and the ways in which a sample can lose its identity even when every tube carries a label. Chapter 5 addresses automated capture of instrument data, the most technically demanding part of most deployments. Chapter 6 treats chain of custody, from receipt through storage, transfer and disposal. Chapter 7 is about validation, done in proportion to risk rather than in proportion to anxiety. Chapter 8 covers the long life of the system after go-live: change control, audit trail review, periodic review, backup and restore, inspection, and eventually retirement. The conclusion draws out what follows from treating all of this as one discipline. A note on scope. The booklet focuses on analytical and quality control laboratories and on GLP test facilities. It touches clinical diagnostic laboratories, which work under different regimes such as ISO 15189 and national clinical laboratory rules, only where the principles are shared. Electronic laboratory notebooks, scientific data management systems and chromatography data systems appear where they meet the LIMS, because in real laboratories the boundaries between these systems are blurred and the evidence has to survive crossing them. A word on regulations and time Regulatory guidance on computerised systems has changed more in the last decade than in the twenty years before it. The FDA finalised its guidance on Computer Software Assurance in September 2025. The European Commission released a substantially expanded draft of Annex 11 for consultation in July 2025, alongside a draft revision of Chapter 4 on documentation and a new draft Annex 22 on artificial intelligence. A revision of the United States Pharmacopeia's general chapter on analytical instrument qualification was proposed in 2025. Readers should always check the current status of the documents cited here before relying on a particular clause. The underlying expectations, however, have been stable for a long time: records must be attributable, legible, contemporaneous, original and accurate; systems must be fit for their intended use and shown to be so; and changes must be controlled and traceable. A laboratory that designs for those expectations will find each new guidance document a refinement rather than a revolution. Chapter 1. What a LIMS Is For: The Reconstruction Test The term Laboratory Information Management System dates from the late 1970s and early 1980s, when minicomputers first became affordable for industrial laboratories and vendors began selling software to track the flow of samples through quality control. The early systems were built for high-volume, repetitive testing in petrochemical, pharmaceutical and environmental laboratories, where the same assays were run on thousands of samples and the main problem was logistics: knowing what had arrived, what was waiting, what had been tested and what had failed. Over four decades the systems acquired much more: specification management, stability study scheduling, instrument interfaces, inventory, reagent tracking, training records, electronic signatures, and web and cloud delivery. ASTM International's Standard Guide for Laboratory Informatics, E1578, now describes the LIMS as one component in a larger informatics landscape that includes electronic laboratory notebooks, scientific data management systems, chromatography data systems and laboratory execution systems. Those descriptions tell a laboratory what a LIMS can do. They do not tell it what a LIMS is for, and a laboratory that deploys one without answering that question tends to end up with an expensive sample register. This chapter proposes an answer and explains why academic laboratories in particular need to hold on to it. The reconstruction test The proposal is that a LIMS in a regulated laboratory exists to make one kind of reconstruction possible. Given any reported result, a competent person who was not present should be able to establish from the records alone: 1.which sample the result belongs to, and that the sample is what it is claimed to be; 1.where the sample came from, who has held it, where it has been stored and under what conditions, and what has been done to it; 2.which method, instrument, reagents, standards and software were used, and that each was in a fit state at the time; 3.who performed the work and who reviewed and approved it, and that each was authorised and trained to do so; 4.what the original data were, how the reported value was derived from them, and whether anything was changed, deleted or repeated along the way, with the reasons. Call this the reconstruction test. It is not a new idea. It is a paraphrase of what GLP and GMP regulations have always required of paper records, and of what the data integrity guidance issued since 2015 has required of electronic ones. The Organisation for Economic Co-operation and Development's Principles of Good Laboratory Practice, for example, require that raw data be retained in a way that allows a study to be reconstructed, and national inspectorates use the phrase "study reconstruction" in exactly this sense. What the test adds is a way of deciding, when configuring or validating a system, which functions matter. A function matters in proportion to the part it plays in passing the test. The test also explains why a LIMS can make things worse. When records are on paper, each page is a separate piece of evidence, and its weaknesses are visible: a crossed-out number, a missing signature, a date in a different ink. When records are in a database, the evidence is the combination of the stored values, the audit trail, the user account model, the system configuration at the time, and the validation that shows the system behaved as configured. Every one of those layers can fail invisibly. A result that looks perfectly clean on a LIMS report may rest on a shared login, an audit trail that was switched off during a data migration, or an instrument interface that silently truncated a decimal place. The report carries the authority of the system without the substance. Records, not results The practical consequence is that the laboratory's attention should be on records rather than results. Scientists naturally think of the result as the product: the concentration, the potency, the sterility outcome. For the reconstruction test the product is the complete record from which the result can be re-derived and checked. That includes the things scientists often regard as clutter: the failed injection, the re-prepared standard, the sample that was received warm, the analyst who was not yet signed off on the method and performed the work under supervision. Regulators have been explicit that this clutter is part of the record. The United Kingdom Medicines and Healthcare products Regulatory Agency's GXP Data Integrity Guidance and Definitions (March 2018) defines raw data as original records and documentation, retained in the format in which they were originally generated, or as a true copy, and it warns against treating a printout or a summary as the record when a richer electronic original exists. The FDA's Data Integrity and Compliance With Drug CGMP: Questions and Answers (December 2018) makes a similar point, and addresses directly the practice of discarding or not reporting results that are later deemed invalid. Its position is that data created as part of a CGMP record must be evaluated by the quality unit as part of release decisions, and that a result cannot be excluded simply because the analyst believes it to be wrong. A LIMS configured to store only final results, with the working-out held elsewhere or nowhere, therefore fails the test by design, however polished it is. One configured to store the relationships between samples, preparations, runs, raw files, calculations, reviews and approvals passes it, even if its interface is clumsy. Why academic laboratories find this hard Academic laboratories that move into regulated work bring strengths that industry often envies: deep scientific expertise, a culture of scrutiny and debate, and a willingness to solve problems from first principles. They also bring habits that sit badly with the reconstruction test, and it is worth naming them plainly because they shape every later decision about configuration and validation. The first is the idea that the scientist is the authority. In discovery research, a senior investigator's judgement about which experiments to believe is the whole point. In regulated work, judgement is still essential, but it must be exercised within a documented procedure and leave a trace. An analyst who decides that a result is an outlier and re-runs the sample is exercising judgement; doing it without recording the original result and the reason is a data integrity failure. The LIMS has to be configured so that judgement is recorded rather than prevented, and the people using it have to accept that recording it is not an insult. The second is the idea that data belong to the person who generated them. Research data are often held on personal drives, laptops or cloud storage accounts that move with the researcher. Regulated records belong to the organisation and must remain available for retention periods that can run for many years after the people who generated them have left. OECD GLP, for example, requires the archiving of raw data and specimens for the period specified by the national authority, and the United States GLP regulations at 21 CFR 58.195 set specific retention periods tied to the regulatory submission. A LIMS that depends on files in a postdoctoral researcher's personal folder does not satisfy those requirements, however carefully the folder is named. The third is staff turnover. Academic laboratories run on students and fixed-term researchers. A regulated facility inside a university may see a third of its analysts change within two years. Every departure takes knowledge of how the system is really used, and every arrival brings someone who must be trained, authorised and given an account with the right permissions. The LIMS and the processes around it must make turnover survivable: role-based access that is easy to administer correctly, training records linked to authorisations, and procedures that do not depend on one person's memory. The fourth is funding. Research infrastructure is often bought with capital grants that pay for acquisition but not for years of maintenance, periodic review, upgrades and revalidation. A LIMS is not a piece of equipment that can be bought and left. The largest share of its lifetime cost falls after go-live, and a system that cannot be maintained will drift out of its validated state. Chapter 8 returns to this problem; it belongs at the outset because it should shape the choice of system and the ambition of the configuration. The fifth is scale. A pharmaceutical quality control laboratory might have dozens of analysts and a dedicated informatics team. An academic GMP facility producing cell therapies for early-phase trials may have a handful of staff, each playing several roles. Regulations that assume separate people for performing, reviewing, approving and administering become difficult to apply. They still have to be applied, because the risks they address do not shrink with the team. The answer is careful design of roles and compensating controls, not exemption. What the LIMS is not The reconstruction test also clarifies boundaries. A LIMS is not usually the right place for free-form experimental narrative, which belongs in a notebook, paper or electronic. It is often not the right place for large raw data files such as chromatograms, mass spectra or images, which are better held in the instrument's own data system or a scientific data management system, with the LIMS holding the link and the derived result. It is not a quality management system, although many LIMS include modules for deviations, corrective and preventive actions and training. And it is not, by itself, a data integrity programme. What matters is that the boundaries are deliberate and documented. For each type of record the laboratory should be able to say which system holds the original, which systems hold copies or derived values, how they are linked, and how the link is protected. A common and serious gap is the undocumented seam: a result typed into the LIMS from a printout, a spreadsheet used to perform a calculation between the instrument and the LIMS, a sample transferred between two systems by exporting and re-importing a file. Each seam is a place where the reconstruction can break, and each is invisible unless someone has drawn the map. Drawing that map is the first real task of a LIMS deployment, and it is worth doing before any software is chosen. For each regulated process, the laboratory traces the path of a sample and its data from receipt to report and archive, noting every system, every hand-off, every calculation and every point at which a person makes a decision. The regulatory literature calls this a data flow diagram or a process map. It can be done on a whiteboard. Its value lies in forcing the question, for every step, of what evidence will exist afterwards. A reconstruction worked through It helps to see the test applied to a single result. Consider an academic facility manufacturing an autologous cell therapy for a phase I trial. Among the release tests for each batch is a flow cytometry assay that reports the percentage of viable cells expressing a transgene. The certificate of analysis states 42 per cent against a specification of not less than 20 per cent. An inspector picks this figure and asks the facility to show where it came from. The first question is identity. The facility must show that the tube analysed was drawn from the batch in question, not from the batch processed the day before for a different patient. That means a record of the sampling event, linking a sample identifier to a batch identifier, with the time, the person and the container. It means that the tube carried that identifier from the moment it was drawn, and that the cytometer acquisition file is associated with the same identifier. If the tube was labelled by hand and the analyst typed a sample name into the cytometer software, the link rests on two transcriptions, each a place where a digit can be swapped. The second is custody and condition. The sample may have been held on ice for two hours before staining because the cytometer was in use. The method may permit that, or it may not. If the LIMS records sampling and analysis times, the interval can be checked. If it records only a date, it cannot. The third is method and instrument. The inspector will want the method version, the lot numbers of the antibodies and viability dye, the cytometer's daily performance check for that day, the compensation settings and the gating template. Some of this lives in the LIMS, some in the cytometer software, some in reagent logs. What matters is that each is linked to the result and that the links were made at the time, not reassembled for the inspection. The fourth is people. The analyst must have been trained on the current method version before the date of the analysis, and the reviewer must have been someone other than the analyst with authority to review. If the facility's LIMS holds training records and enforces them at the point of result entry, this is quick to show. If training records are in a separate spreadsheet, it can still be shown, but with more effort and less certainty. The fifth is derivation. The figure of 42 per cent is the output of gating decisions applied to an acquisition file. The inspector will ask whether the gates were adjusted, whether the file was re-analysed, whether a first acquisition was discarded, and why. The cytometer software's own audit trail and the LIMS audit trail together should answer that. If the analyst exported a statistics table, pasted it into a spreadsheet and typed the percentage into the LIMS, the chain has three seams, and the spreadsheet may not have an audit trail at all. Nothing in this example requires exotic technology. It requires that the system was designed around the questions the inspector would ask, and that someone checked, before the batch was made, that each question had an answer in the records. That is the discipline the rest of this booklet describes. The shape of a deployment With the test in view, a LIMS deployment has a recognisable shape. It begins with the map just described and with a statement of intended use: which regulated processes the system will support, under which regulations, and with what boundaries. It moves to configuration, in which the laboratory's sample types, tests, specifications, workflows and roles are expressed in the system's terms. It proceeds to identity and custody, where the physical world of tubes, plates, freezers and couriers is tied to the records. It confronts instrument integration, where the choice between automated capture and manual entry has the largest effect on data integrity. It is validated, in proportion to risk, so that the laboratory can show that the configured system does what it needs to do. And then it enters a long period of operation in which the validated state must be maintained against continuous change. None of these stages is purely technical. Each depends on decisions about how the laboratory works and what it is prepared to record, and each is easier if the people making those decisions keep asking the same question. If an independent person arrived tomorrow and picked a result at random, could they tell its whole story from what the system holds? The remaining chapters take each stage in turn and show how that question shapes the answer. Chapter 2. The Regulatory Ground A laboratory preparing to deploy a LIMS in a regulated setting soon discovers that there is no single rulebook. There are statutes and regulations that carry legal force, guidance documents that describe how regulators interpret them, international principles adopted into national law, pharmacopoeial chapters, industry good practice guides, and standards bodies' documents. They overlap, use different vocabularies, and were written at different times for different problems. The purpose of this chapter is not to summarise all of them, which would take a much longer book, but to show how they fit together around the reconstruction test and to identify the few ideas that recur in all of them. Three regimes and what separates them Regulated laboratory work in academic institutions usually falls under one of three regimes, and the first task is to know which. Good Laboratory Practice applies to non-clinical safety studies whose results will be submitted to regulators: toxicology, safety pharmacology, some ecotoxicology and environmental fate studies, and certain analytical work supporting them. Its origins lie in the scandals of the 1970s, when FDA investigations found fabricated and poorly controlled safety data at contract laboratories, most notoriously Industrial Bio-Test Laboratories. The FDA's GLP regulations, 21 CFR Part 58, were published as a final rule in December 1978. The OECD adopted its Principles of GLP in 1981 alongside a Council Decision on the Mutual Acceptance of Data, under which studies performed to the principles in one member country are accepted by others; the principles were revised in 1997. In the European Union they are implemented through Directive 2004/10/EC and national monitoring authorities, such as the MHRA's GLP Monitoring Authority in the United Kingdom. GLP is organised around the study: a study director carries single-point responsibility, a protocol defines the work, raw data and specimens are archived, and an independent quality assurance unit inspects. The reconstruction test is native to GLP. Good Manufacturing Practice applies to the manufacture and quality control of medicines, including investigational medicines for clinical trials. For an academic laboratory this usually means a facility producing cell or gene therapy products, radiopharmaceuticals, or small-batch investigational products, together with the quality control laboratory that tests them. In the United States the core regulations are 21 CFR Parts 210 and 211, with 21 CFR 211.68 addressing automatic, mechanical and electronic equipment and 21 CFR 211.194 addressing laboratory records. The FDA's 2008 guidance on CGMP for Phase 1 Investigational Drugs describes a proportionate approach for the earliest clinical stages, and cell and tissue products also fall under 21 CFR Part 1271. In the European Union the principles are set out in EudraLex Volume 4, with Annex 11 addressing computerised systems, Annex 13 investigational products, and Part IV the specific GMP requirements for advanced therapy medicinal products, published in 2017. GMP is organised around the batch: a quality unit independent of production approves materials and releases product, and every batch record must show that the product meets its specification and was made under control. The third regime is accredited testing, most often under ISO/IEC 17025:2017, which sets general requirements for the competence of testing and calibration laboratories. Clause 7.11 addresses the control of data and information management, including validation of laboratory information management systems before use. ISO 20387:2018 does similar work for biobanks, and ISO 15189 for medical laboratories. These are not regulations in the legal sense, but accreditation is often a contractual or regulatory condition, and assessors look at LIMS configuration and validation in much the same way inspectors do. The regimes differ in important ways. GLP centres on the study and its reconstruction years later; GMP centres on the batch and the release decision; ISO 17025 centres on the competence of the laboratory and the validity of its results. A LIMS may serve all three in the same institution, and a facility running both a GLP analytical service and a GMP quality control laboratory will need to decide whether to configure one system for both or keep them apart. Either can work. What does not work is leaving the question unanswered, so that nobody can say which regulations govern a given record. Electronic records and signatures Layered over the three regimes are rules specifically about electronic records. In the United States the governing regulation is 21 CFR Part 11, Electronic Records; Electronic Signatures, published in March 1997 and effective in August of that year. It applies to records required by any FDA predicate rule, such as Part 58 or Part 211, when they are kept in electronic form, and to electronic signatures used in place of handwritten ones. Its requirements for closed systems include validation, the ability to generate accurate and complete copies, protection of records for their retention period, limited system access, secure computer-generated time-stamped audit trails that record the creation, modification and deletion of records without obscuring previous values, operational and authority checks, device checks, training, and controls over system documentation. Its signature provisions require that each signature be unique to one individual, that its meaning be recorded, and that it be linked to the record so that it cannot be copied or transferred. The first years of Part 11 were difficult. Industry found the requirements broad and expensive, and the FDA withdrew its early guidance. In August 2003 it issued Part 11, Electronic Records; Electronic Signatures: Scope and Application, which narrowed its interpretation and announced enforcement discretion on validation, audit trails, record retention and copying for certain records, while making clear that predicate rule requirements still applied in full. That guidance remains in force and still shapes how Part 11 is read. Its central message, that the requirements should be applied on the basis of a justified and documented risk assessment, anticipated much of what followed. In the European Union the equivalent is Annex 11 to the GMP guide, Computerised Systems, in its 2011 revision. It is short, about five pages, and principle-based. It requires risk management across the life cycle, defined responsibilities for process owners and system owners, supplier assessment, validation proportionate to risk, data checks for manual entry and interfaces, accuracy checks, data storage and backup, printouts, audit trails based on risk, change and configuration management, periodic evaluation, security, incident management, electronic signatures, batch release controls, business continuity and archiving. For GLP, the OECD's Advisory Document No. 17, Application of GLP Principles to Computerised Systems (2016), closely follows the structure of Annex 11, and a supplement on GLP and cloud computing followed in 2023. In July 2025 the European Commission and the European Medicines Agency released a draft revision of Annex 11 for public consultation, which closed in October 2025. The draft is far longer and more prescriptive than the text it would replace. It gives detailed attention to audit trails, including the recording of old and new values and the reasons for change and the review of audit trails before batch release; to identity and access management, including unique accounts, restrictions on shared accounts and multi-factor authentication for remote access to critical systems; to periodic review as a structured exercise; and to contracts with service providers, including cloud providers. At the time of writing, laboratories should check whether the final text has been adopted and from what date it applies. Even as a draft it is a clear statement of what European inspectors expect. Data integrity: the organising idea Between about 2013 and 2021 a series of high-profile enforcement actions, many involving quality control laboratories, turned regulatory attention from the formal requirements of Part 11 and Annex 11 to a broader question: are the data true? Inspectors found analysts running unofficial trial injections before the official ones, deleting or overwriting failing results, disabling audit trails, sharing administrator accounts, and backdating records. None of these required sophisticated fraud; most exploited weaknesses in how systems were configured and administered. An earlier case, at Able Laboratories in 2005, where an FDA inspection found that analysts had substituted and manipulated chromatographic data and the company recalled its products and ceased operations, had already shown what was at stake. The response was a wave of guidance: the MHRA's GXP Data Integrity Guidance and Definitions in its March 2018 revision, the FDA's Data Integrity and Compliance With Drug CGMP Questions and Answers in December 2018, the World Health Organization's guideline on data integrity in 2021, the Pharmaceutical Inspection Co-operation Scheme's PI 041-1, Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments, effective July 2021, and the OECD's Advisory Document No. 22 on GLP Data Integrity in 2021. Industry added ISPE's GAMP Guide: Records and Data Integrity in 2017. All of them organise their expectations around the acronym ALCOA, usually credited to Stan Woollen of the FDA in the 1990s: data should be Attributable, Legible, Contemporaneous, Original and Accurate. Later formulations, often written ALCOA+, add Complete, Consistent, Enduring and Available. These are not abstractions. Each maps directly to LIMS functions. Attributable means unique user accounts and audit trails that identify who did what. Legible means records that can be read, and understood, for their whole retention period, including the metadata that give them meaning. Contemporaneous means recording at the time of the activity, which in a LIMS means server-controlled time stamps and workflows that do not allow results to be batched up and entered at the end of the week. Original means the first capture of the data, or a verified true copy, which is why the question of whether instrument data are captured automatically or transcribed matters so much. Accurate means correct, which depends on validation, on checks at data entry, and on calculations that have been verified. Complete means that nothing is missing, including failed runs and repeats. Consistent means that the sequence of events is coherent. Enduring and Available mean that records survive and can be retrieved for as long as they must be kept. The guidance also introduced or sharpened several ideas that shape LIMS deployments directly. One is data governance: the organisational arrangements, including culture, that ensure data integrity, as distinct from the technical controls. Another is the distinction between static records, such as a paper printout or a fixed image, and dynamic records, such as a chromatogram that can be reprocessed or a LIMS record whose meaning depends on its links to other records. The MHRA guidance is explicit that a static printout of a dynamic electronic record is not an adequate substitute for the original. A third is the expectation of risk-based review of audit trails as part of routine data review, not only during investigations. A fourth is the idea of data criticality and data risk: the effort spent controlling a record should reflect its effect on product quality, patient safety or study conclusions, and how vulnerable it is to unauthorised change. How the frameworks fit together For a laboratory reading these documents for the first time, the overlap can be bewildering. The key frameworks and what each asks of a LIMS are summarised in Table 1. Table 1. Principal frameworks governing LIMS in regulated laboratories. Framework Issuer and date Scope What it asks of a LIMS 21 CFR Part 11 FDA, 1997; scope guidance 2003 Electronic records and signatures under FDA predicate rules Validation, audit trails, access control, signature linking 21 CFR Part 58 FDA, 1978 Non-clinical laboratory studies (GLP) Raw data retention, study reconstruction, archiving 21 CFR Parts 210 and 211 FDA Drug manufacture and quality control (GMP) Equipment checks (211.68), complete laboratory records (211.194) EU GMP Annex 11 European Commission, 2011; draft revision 2025 Computerised systems in GMP Risk-based validation, supplier assessment, periodic review, audit trails OECD Advisory Documents No. 17 and No. 22 OECD, 2016 and 2021 Computerised systems and data integrity in GLP Life-cycle validation, cloud controls, raw data definition MHRA GXP Data Integrity Guidance MHRA, 2018 All GxP data ALCOA+, data governance, true copies, audit trail review PIC/S PI 041-1 PIC/S, 2021 GMP and GDP data management Detailed data integrity controls for computerised systems The table conceals as much as it shows, because the frameworks are not alternatives. A GMP quality control laboratory in the European Union that tests investigational products for a trial with United States sites may need to satisfy Annex 11, Part 11 and the relevant parts of Part 211 at once, and will be inspected against PIC/S PI 041 thinking regardless. The practical approach is to identify the most demanding applicable requirement on each point, design to it, and record in the validation plan which regulations were considered. Industry guidance and the move to risk-based assurance Regulations say what must be achieved; they say little about how. The gap has been filled mainly by the International Society for Pharmaceutical Engineering's GAMP guidance. GAMP 5, A Risk-Based Approach to Compliant GxP Computerized Systems, first published in 2008 and revised as a second edition in July 2022, is the most widely used framework for computerised system validation. It classifies software into categories: infrastructure software (category 1), non-configured products (category 3), configured products (category 4) and custom applications (category 5). Most commercial LIMS are category 4, with category 5 elements wherever custom code, scripts or interfaces are added. GAMP scales the validation effort to the category and to the risk of each function. The second edition put particular emphasis on critical thinking, on leveraging supplier activities rather than duplicating them, and on modern practices such as agile development, cloud services and automated testing. The same direction of travel is visible at the FDA. Its guidance on Computer Software Assurance for Production and Quality System Software, issued in draft in September 2022 and finalised in September 2025, formally addresses software used in medical device production and quality systems. Its underlying message has been taken up across the life sciences: assurance activities should focus on software functions that could affect product quality or patient safety, testing should be proportionate, unscripted and exploratory testing are acceptable where risk is lower, and documentation should record what was done and found rather than generate paper for its own sake. For an academic laboratory, this is good news. The idea that validation means thousands of pages of scripted tests for every screen of the system was always a misreading, and it is now explicitly rejected by regulators and industry guidance alike. The laboratory is expected to think: to understand which functions of its LIMS carry the weight of the reconstruction test, to test those thoroughly, to rely on supplier evidence where it is sound, and to explain its reasoning. Chapter 7 develops this in detail. Academic particularities A few regulatory points bear particularly on academic institutions. First, the regulations apply to the regulated activity, not to the institution. A university is not exempt from GMP because it is a university. If it manufactures an investigational medicinal product, the facility needs the appropriate authorisation, and its quality control laboratory must meet GMP. If it conducts a study intended to support a regulatory submission and claims GLP compliance, it must be in the national GLP monitoring programme where one exists and must meet the principles in full. Second, the boundary between regulated and unregulated work inside the same institution must be clear. A research group may use the same flow cytometer for GMP release testing and exploratory research. The LIMS configuration, user roles and data locations must keep the two apart, or the regulated controls must apply to both. Third, the institution's central information technology services are usually not accustomed to GxP requirements. They may apply operating system patches automatically, restore backups without documentation, or move servers between data centres without notice. Each of those is a change to a validated system. The facility needs a written agreement with its IT provider that defines responsibilities, and the provider's staff need enough training to understand why the agreement exists. The same applies with greater force to cloud and software-as-a-service providers, where the laboratory has less visibility and must rely on contracts and supplier assessment. These points lead naturally into configuration, which is where the regulatory expectations are turned into the concrete structure of the system the laboratory will use every day. Chapter 3. Configuring the System Around the Evidence Most commercial LIMS are sold as configurable products. Out of the box they contain a general model of laboratory work: samples, tests, analyses, specifications, instruments, users and reports, connected by workflows. The laboratory's job is to express its own processes in those terms. This is configuration, and it is where most of the decisions that determine whether the system will pass the reconstruction test are made. It is also where academic deployments most often go wrong, usually not through technical error but through configuring the system to mirror what people currently do rather than what the records need to show. Configuration, customisation and why the distinction matters A first distinction is between configuration and customisation. Configuration means setting up the product using the tools the vendor supplies for that purpose: defining sample types, building test templates, setting up specification limits, arranging workflow states, assigning permissions. Customisation means changing the product's behaviour with code: scripts, stored procedures, custom screens, bespoke interfaces or modifications to the vendor's source. In GAMP terms, configuration keeps the system in category 4; customisation adds category 5 components. The distinction matters for three reasons. First, custom code carries more risk because it has not been tested by the vendor or by other customers, so it requires more validation effort. Second, customisations are fragile across upgrades. A script that works in version 10 of a product may break silently in version 11, and every upgrade then requires regression testing of every customisation. Third, customisations accumulate knowledge in the heads of the people who wrote them. In an academic setting, where the person who wrote the script may be a doctoral student who has since left, that knowledge is easily lost. The practical rule is to configure wherever possible and customise only where configuration genuinely cannot meet a requirement that matters for the reconstruction test. When customisation is necessary, it should be specified, reviewed, version-controlled, tested and documented as software, not as a quick fix. Many laboratories find, on reflection, that a requirement they thought needed custom code was really a preference about how a screen should look, and that the standard product's way of doing things was acceptable once people had used it for a month. The sample model The heart of any LIMS configuration is the sample model: how the system represents the things being tested and the relationships between them. Getting it right is more important than any other configuration decision, because almost every other record hangs from it, and because it is very difficult to change once data have accumulated. A sample model answers several questions. What is a sample? In a GMP quality control laboratory, the answer might be a container drawn from a batch at a defined point in the process. In a GLP toxicology study, it might be a blood sample from a particular animal at a particular time point, or a tissue taken at necropsy. In a biobank, it might be a donation that is later divided into many aliquots. The model must distinguish the physical container from the material it contains and from the logical entity that tests are requested against. How are samples related? Most real laboratories need parent and child relationships: a primary sample divided into aliquots, a tissue homogenised into a lysate, a batch sampled at several stages, a pooled sample drawn from several sources. The LIMS must record these relationships so that a result on a child can be traced to its parent and, through it, to the batch, study or donor. It must also record derivations that change the nature of the material: extraction, dilution, digestion, staining. A result on a diluted extract is meaningless without the dilution factor and the link back to the original. What metadata must travel with a sample? For a batch sample, the batch number, the sampling point, the time of sampling, the container type and the storage condition. For a study sample, the study number, the animal or subject identifier, the dose group, the time point and the matrix. For a donor sample, consent status and any restrictions on use. The temptation is to capture everything anyone might ever want. The discipline is to capture what the reconstruction test needs and what downstream decisions depend on, and to make those fields mandatory and validated rather than optional free text. A hypothetical but entirely typical case shows why this matters. Suppose an academic toxicology facility configures its LIMS with a single sample type for plasma, using free-text fields for the study number and time point. Over two years, analysts enter time points as "2h", "2 h", "120 min", "T2" and "2hr". Reports that group results by time point silently split what should be single groups, and a pharmacokinetic analysis for a study destined for a regulator has to be redone by hand once the problem is noticed. The fix is not complicated: a controlled list of time points defined per study in the protocol and enforced at sample registration. But making it retrospectively requires a documented data correction across thousands of records, each with a reason in the audit trail. A day spent on the sample model at the outset avoids months of remediation. Anyone who has worked in laboratory informatics will recognise the pattern. Tests, methods and specifications Next comes the representation of what is done to samples. A LIMS typically distinguishes a test or analysis, meaning a defined procedure producing one or more results, from the method document that describes it, and from the specification that defines what results are acceptable. Each test template should carry a version that is tied to the approved method document. When the method changes, the template changes, and results should record which version was used. Many systems allow templates to be edited in place; in a regulated setting that is dangerous, because it changes the meaning of historical records. Templates should be versioned, with old versions retired rather than overwritten. Each result should carry its units, its reportable precision and its rounding rule, and these should be configured rather than left to the analyst. A result of 0.0495 reported against a limit of not more than 0.05 passes or fails depending on whether it is rounded before or after comparison, and on how many decimal places are retained. Pharmacopoeias and regulatory guidance give rules for this. The LIMS should implement them consistently, and validation should test the boundary cases. Calculations deserve particular care. A LIMS that calculates a result from several entered values, for example a concentration from peak areas, weights, dilution factors and a standard's purity, is performing a function with direct evidential weight. The formula must be verified during validation, and the entered values and the calculated result must both be retained. Where calculations are done outside the LIMS, in instrument software or a spreadsheet, the laboratory must decide where the validated calculation lives and ensure that it is not repeated or overridden elsewhere. Specifications should be held in the system as controlled master data, with an effective date and an approval record, so that each result is compared against the specification in force at the time. The system should flag out-of-specification results automatically and route them into the laboratory's investigation procedure. It should not allow them to be quietly re-entered as passing values. Regulators have consistently criticised laboratories that invalidated failing results without investigation; the LIMS should make that difficult by design. Workflows and states A LIMS workflow defines the states a sample or result passes through and who may move it between them. A typical sequence for a result runs from pending, through entered, to reviewed, approved and reported, with branches for rejected, retested or cancelled. Each transition is an event in the audit trail and, for review and approval, often an electronic signature. The design principle is that workflows should encode the laboratory's procedures, not bypass them. If the procedure says that every result is reviewed by a second person before release, the workflow must require a review step and prevent the reviewer from being the person who entered the result. If the procedure allows a result to be corrected after review, the workflow must send it back for re-review and record the reason. If a sample may be cancelled, cancellation must require a reason and should itself be reviewable. Two common faults are worth watching for. One is the superuser shortcut: an account with permission to move records directly to any state, used during commissioning to fix problems and never removed. The other is the silent reopening, where an approved result can be edited without returning to the review state, so that the approval signature ends up attached to a value it never saw. Both defeat the reconstruction test and both are easily found in an inspection. Roles, permissions and the small-team problem Every regulatory framework expects access to be limited to authorised individuals and actions to be attributable to them. In a LIMS this is implemented through user accounts and roles. Each person has a unique account, never shared. Each account is assigned one or more roles, and each role carries permissions: to register samples, to enter results, to review, to approve, to edit master data, to administer the system. The regulations also expect segregation of duties. Those who enter results should not approve them. Those who configure the system should not use it to generate regulated data. And, most important, those who can change audit trail settings, delete records, alter system time or grant permissions should be independent of those whose work the controls protect. The draft revision of Annex 11 is explicit on this last point. It expects administrative privileges to be held by people without a conflict of interest in the data. In a large laboratory this is straightforward. In an academic facility with five staff it is not. The facility manager may also be the most experienced analyst, the head of quality may be part-time, and the only person who understands the LIMS configuration may be a scientist who also runs assays. The regulations do not make an exception for small teams, but they do allow a thoughtful approach. The usual solutions combine several measures. First, separate the administrative function from the scientific one wherever possible. A system administrator in the university's research computing service or quality department, who has no role in generating results, can hold the privileged account. Scientists who need to maintain master data, such as test templates and specifications, can be given a configuration role that allows controlled changes but not access to audit trail settings or user management. Second, where one person must hold incompatible roles, give them separate accounts for each and use them for separate purposes, with every use of the privileged account logged and reviewed by someone else. This is a compensating control, not an ideal, and it should be documented as such in the risk assessment. Third, make review of privileged activity a routine task. A monthly review of the administrator account's audit trail by the head of quality, recorded and signed, turns an unavoidable concentration of power into a supervised one. Fourth, link roles to training. Many LIMS can hold training records and prevent a user from entering results for a method they are not trained on. In an academic facility where staff turn over frequently, this is one of the most valuable configuration choices available, because it makes the question of whether an analyst was authorised on a given day answerable from the system rather than from memory. Finally, manage the life cycle of accounts. Leavers should be disabled promptly; accounts should never be deleted, because historical records must remain attributable. Periodic access reviews should confirm that every active account belongs to a current member of staff with a current need for its permissions. Academic laboratories are particularly prone to orphan accounts belonging to visiting researchers, summer students and staff who moved to other groups. Reagents, standards and equipment as linked records A result depends not only on the sample and the method but on the materials and equipment used to produce it. Reference standards, reagents, media, antibodies, calibrators and consumables each have identities, lot numbers, expiry dates and sometimes assigned values such as purity or potency. Instruments and equipment have qualification status, calibration due dates and maintenance histories. The reconstruction test requires that each of these can be linked to each result. Most LIMS offer inventory and equipment modules for this purpose, and they are worth configuring even if the laboratory's first instinct is to leave them for a later phase. The value lies in enforcement at the point of use. If a reference standard is registered in the LIMS with its expiry date and certificate value, and the test template requires the analyst to select the standard lot used, the system can refuse an expired lot and can pull the certificate value into the calculation rather than relying on the analyst to type it. If a balance is registered with its calibration due date, the system can block its selection when calibration has lapsed. Each of these checks prevents a class of error that is common, tedious to detect in review, and embarrassing when found by an inspector. The same logic applies to prepared solutions. A mobile phase, buffer or working standard made in the laboratory is a derived material with its own identity, preparation record, components and expiry. Recording preparations in the LIMS, with their component lots and the person who made them, lets the laboratory answer questions such as which results used a buffer later found to have been made with the wrong salt. Without that link, the answer depends on searching notebooks. There is a cost. Every link that the system enforces is a step the analyst must perform, and a system that demands too many selections for every result will be resented and, in time, circumvented. The laboratory should decide, method by method, which materials and equipment have enough influence on the result to justify enforced linkage, and record the others in simpler ways. That judgement is itself part of the risk assessment that Chapter 7 describes. Configuring in stages A final practical question is how much to configure before going live. Vendors and project managers often favour a comprehensive first release; academic facilities often lack the staff to specify, test and train on everything at once. A staged approach usually works better: begin with the processes whose records carry the most weight, such as release testing for the facility's main product or the analytical phase of its most important GLP study type, configure and validate them thoroughly, and add further sample types, methods and modules in controlled releases. Each release is a change to a validated system and follows the same discipline, but the laboratory learns from each one, and the sample model is tested against real use before it has to carry the whole operation. Master data and configuration as controlled records Everything described in this chapter, from sample types to specifications to roles, is configuration, and configuration is itself a regulated record. It determines how the system behaves and therefore what the records mean. A change to a specification limit or a calculation formula changes the interpretation of every result that passes through it afterwards. Configuration should therefore be managed under change control, with a record of what was changed, why, by whom, when, who approved it and what testing was done. Many LIMS keep an audit trail of master data changes; the laboratory should check that it does and that the trail is reviewed. It is good practice to maintain a configuration specification, a document or export that describes the configured state of the system at each version, so that the configuration in force on any date can be established. For an inspector reconstructing a result from three years ago, knowing which specification version and which calculation were active on that date is as important as knowing the result itself. A final configuration principle ties the chapter together. For each configuration choice, ask what evidence it will generate and whether that evidence will answer the questions an independent reviewer would ask. A sample model that records parentage answers the question of where an aliquot came from. A versioned test template answers the question of which method was used. A workflow that separates entry from approval answers the question of whether the result was checked. A role model with separated administration answers the question of whether anyone could have changed the record without a trace. Configuration that is designed this way is not merely compliant; it is useful, because the same evidence that satisfies an inspector also lets the laboratory find and fix its own mistakes. Hashtags: #LaboratoryInformationManagementSystems #LIMS #LaboratoryInformatics #RegulatedLaboratories #GMPCompliance #GLPCompliance #DataIntegrity #ALCOAPlus #ElectronicRecords #AuditTrails #ChainOfCustody #SampleManagement #InstrumentIntegration #ComputerSystemValidation #RiskBasedValidation #GAMP5 #CFRPart11 #EUAnnex11 #LaboratoryCompliance #QualityControl #ChangeControl #AccessControl #DataGovernance #RegulatoryInspection #FutureOfLaboratoryInformatics
- The Peer Review Ecosystem (Ethics, Editorial Roles, and Constructive Critique)
Download the Book (PDF): Introduction The first review most scientists write arrives without ceremony. An email from an editor they have never met, a manuscript title that sits somewhere near their own work, a deadline two or three weeks away, and a link to a submission system with a text box and a drop-down menu of recommendations. There is rarely any training attached. The graduate student or postdoctoral researcher who clicks "accept invitation" is expected to know, somehow, what the editor wants, what the authors deserve, what counts as a fatal flaw and what counts as a matter of taste, and how to say all of it in prose that will be read by strangers who have spent years on the work in question. Most early-career reviewers fill this gap by imitation. They recall the reviews they have received themselves, the sharp ones that stung and the lazy ones that missed the point, and they write something in between. Some become harsh, on the theory that rigour and severity are the same thing. Some become timid, afraid that a junior person has no standing to criticise a senior laboratory. Many produce long lists of small objections, because small objections are easy to find and feel like diligence. Very few are ever told whether their reviews were any good. This book is written for that reviewer. It treats peer review not as a single act of judgement but as an ecosystem: a set of roles, each with distinct obligations, that together produce the decisions which shape what enters the scientific record. Authors, reviewers, handling editors, editors-in-chief, publishers, research integrity officers, readers who comment after publication, and increasingly the platforms that publish reviews openly all occupy positions in this system. A reviewer who understands only their own position will misjudge what their report is for. A reviewer who understands the whole system can write a report that does its job. The argument The controlling claim of this book is simple to state and harder to practise. The ethics of peer review and the craft of peer review are the same subject. Fairness, confidentiality, disclosure of conflicts, and honesty about the limits of one's own expertise are usually taught, when they are taught at all, as a compliance layer: rules to observe before the real intellectual work begins. That framing is wrong. A review distorted by an undeclared rivalry is not an ethical failure sitting beside a technically excellent assessment; it is a bad assessment. A review that hides its uncertainty behind confident language misleads the editor about the evidence. A review that is contemptuous in tone gets its substantive points ignored, which means the flaws it found are more likely to survive into print. Constructive critique is not critique with the edges sanded off. It is critique that is accurate about the work, accurate about the reviewer's own position, and aimed at a decision someone else must make. Three commitments follow from that claim, and they run through every chapter. The first is that a review is written for two audiences with different needs. The editor needs help making a decision: is this work sound, is it important enough for this venue, and what would it take to fix? The authors need help improving the work, whatever the decision. A report that serves only one of these audiences is half a report. The second is that the reviewer is a witness, not a judge. Reviewers advise; editors decide. This matters practically, because it changes how a recommendation should be framed, and ethically, because it limits what a reviewer is entitled to do when something looks wrong. A reviewer who suspects manipulated data is not the investigator, the prosecutor, or the jury. They are the person who noticed, and their obligation is to report what they noticed clearly and to the right person. The third is that peer review is a human process with known and measurable failure modes. Reviewers disagree with each other far more than most scientists assume. They miss errors deliberately planted in manuscripts. They are swayed by the prestige of authors and institutions. They are, at times, cruel. None of this is a reason for cynicism, but all of it is a reason for method. A reviewer who knows where the process tends to fail can build habits that guard against those failures in their own work. What the chapters do The first two chapters establish the system. Chapter 1 asks what peer review is for, tracing how a practice that most scientists assume is ancient became standard only in the second half of the twentieth century, and summarising what the evidence says about its reliability. Chapter 2 maps the editorial system: who does what between submission and decision, what editors actually need from reviewers, and how a reviewer's report is used once it leaves their hands. Chapters 3 and 4 are about the core craft. Chapter 3 sets out a method for reading a manuscript critically, from the first pass that establishes what the paper claims to the detailed interrogation of design, analysis, and reporting. Chapter 4 turns to writing: how to structure a report, how to separate essential problems from preferences, how to phrase criticism so that it is heard, and how to make a recommendation that respects the editor's role. Chapters 5, 6, and 7 address the ethical terrain directly, though each treats ethics as part of accurate assessment. Chapter 5 covers conflicts of interest and confidentiality, including the newer question of whether a reviewer may feed a confidential manuscript to a generative AI tool. Chapter 6 examines implicit bias: what the experimental evidence shows about status, gender, and confirmation effects in review, and what an individual reviewer can do about biases they cannot see directly. Chapter 7 deals with disputes over data, from suspected image manipulation to authors who refuse to share the data behind their claims, and sets out what a reviewer should and should not do when something looks wrong. Chapter 8 looks at the changing shape of the system itself. Open identities, published reports, reviewed preprints, and portable reviews that travel between journals are no longer experiments at the margins. Several prominent journals now publish reviewer reports routinely, and at least one major life-sciences journal has abandoned accept-or-reject decisions after review altogether. These models change what a review is, who reads it, and what it is worth to the person who wrote it. The conclusion draws these threads into an argument about what an early-career reviewer should actually do differently, and about which problems in the system remain open. What this book leaves out The book concentrates on manuscript review for journals and preprint review platforms, because that is where most early-career scientists begin. Grant review shares many of the same ethical principles, and the discussion of conflicts, bias, and confidentiality applies there with little change, but the mechanics of study sections and funding panels differ enough that they are treated only in passing. The book also does not attempt to cover every discipline's conventions. Its examples come mostly from the life, physical, and social sciences, where the journal article remains the main unit of communication. Reviewers in fields where conference proceedings dominate, such as much of computer science, will find the principles transferable, and some of the most instructive evidence about bias in review comes from exactly those venues. Finally, the book does not offer templates to be filled in. Checklists have a place, and several appear in the chapters where they help, but a review assembled from a checklist without judgement reads like one. The aim is to give a new reviewer a clear enough picture of the system, and of their own role within it, that they can make good judgements in situations no checklist anticipated. A note on standing Early-career scientists often ask whether they have the standing to review at all, particularly when the authors are senior. The answer is that standing in peer review comes from expertise in the specific question at hand, not from seniority. An editor who invites a postdoctoral researcher usually does so because that person has recently worked with the precise technique, dataset, or model system the manuscript depends on, and often knows its pitfalls better than anyone else in the field. The appropriate response to that invitation is neither deference nor bravado but candour: say what you can assess with confidence, say what lies outside your competence, and do the work carefully. That candour is the thread that connects everything that follows. It is what makes a review useful to an editor, fair to authors, and worth the hours it takes to write. Chapter 1. What Peer Review Is For Ask a room of scientists when peer review began and many will point to the seventeenth century. In 1665 Henry Oldenburg, secretary of the Royal Society of London, launched the Philosophical Transactions, and the journal is often described as the origin of the practice. The description is misleading in an instructive way. Oldenburg edited the Transactions as a private venture, and he selected material largely on his own judgement and through his correspondence. The Royal Society did not take formal responsibility for the journal until 1752, and the system of sending papers to members for written reports developed gradually through the nineteenth century. Even then, it was one practice among several. Many journals were run by editors who decided alone or with a small circle of trusted advisers. The historian Melinda Baldwin has shown how recent the modern expectation really is. Nature, for most of its history after its founding in 1869, relied heavily on editorial judgement, and it did not require external refereeing for all research papers until the 1970s. Across much of science, universal external review became standard only in the decades after the Second World War, driven by the explosive growth of research funding, the resulting flood of submissions, and a growing need for scientific institutions to demonstrate accountability to the governments that paid for them. Peer review, in Baldwin's account, became a marker of scientific legitimacy in public life at roughly the same time it became routine inside journals. A well-known episode from 1936 captures the transition. Albert Einstein and Nathan Rosen submitted a paper on gravitational waves to Physical Review, arguing that such waves might not exist. The editor, John Tate, sent it to a referee, whose anonymous report identified a serious error. Einstein, accustomed to the German journals where editors published the work of eminent authors without external review, was indignant that the manuscript had been shown to a colleague before publication. He withdrew it and published elsewhere. The historian Daniel Kennefick later established that the referee was the cosmologist Howard Percy Robertson, and that Robertson subsequently helped Einstein's assistant understand the problem. The published version, when it appeared in the Journal of the Franklin Institute, reached a quite different conclusion from the original. The referee had been right. The story is often told as a joke at Einstein's expense. It is better read as a reminder that the practice now taken for granted was contested within living memory, and that its authority rests less on tradition than on whether it actually improves the work that passes through it. The functions a review serves Peer review is asked to do several jobs at once, and much confusion about how to review well comes from failing to separate them. At least four can be distinguished. The first is quality control: checking that the methods are sound, the analyses appropriate, and the conclusions supported by the evidence. This is the function most scientists have in mind when they speak of peer review, and it is the one a reviewer is best placed to perform, because it draws directly on technical expertise. The second is selection: judging whether the work is important, novel, or interesting enough for the particular venue. A methodologically flawless study may still be a poor fit for a journal that publishes only work of broad significance. Selection judgements are more subjective than quality judgements, more dependent on the journal's editorial policy, and more vulnerable to bias. Many journals now ask reviewers to separate the two explicitly, and some, notably PLOS ONE since its launch in 2006 and a number of other "soundness-only" journals since, have removed the importance criterion altogether, asking reviewers to judge whether the work is rigorous rather than whether it is exciting. The third is improvement. Reviews routinely make manuscripts better: they catch errors, suggest additional controls or analyses, point to missing literature, and force authors to state their claims more precisely. Surveys of authors consistently find that most believe their published papers were improved by review, even when they resented the process. The improvement function operates regardless of the decision. A rejected paper, revised in light of thoughtful reports, often goes on to a better version at another journal. The fourth is certification. Publication in a peer-reviewed venue signals to readers, hiring committees, funders, journalists, and courts that the work has passed some form of expert scrutiny. This is the function that gives peer review its public weight, and it is also the one most often overstated. Review certifies that a small number of experts, working with limited time and without access to the raw data, found no fatal problem with the manuscript as presented. It does not certify that the findings are true. A reviewer's report can serve all four functions, but not equally. It serves quality control and improvement best when it is specific and technical. It serves selection best when the reviewer states their view of the work's significance separately from their view of its soundness, so that the editor can weigh each against the journal's own priorities. And it serves certification best when it is honest about what the reviewer could and could not check. What the evidence says about reliability For a practice so central to science, peer review was subjected to systematic study remarkably late. The International Congress on Peer Review and Scientific Publication, first held in Chicago in 1989 under the sponsorship of JAMA, began to change that, and a substantial body of research now exists. Its findings are sobering, and an early-career reviewer should know them. Reviewers agree with each other far less than intuition suggests. A 2010 meta-analysis by Lutz Bornmann, Rüdiger Mutz, and Hans-Dieter Daniel, published in PLOS ONE, pooled studies of inter-reviewer agreement on journal manuscripts and found levels of agreement that were low by any conventional standard: reviewers assessing the same manuscript agreed only modestly better than would be expected by chance. Earlier, Peter Rothwell and Christopher Martyn had examined reviews of submissions to two clinical neuroscience journals and conference abstracts, reporting in Brain in 2000 that agreement between reviewers on whether to publish was little better than chance. Grant review shows similar patterns. Low agreement is not in itself proof that review is broken. Editors deliberately choose reviewers with different expertise, and two reviewers who assess different aspects of a paper will naturally reach different overall views. But low agreement does mean that any single review is a noisy signal, and that the recommendation at the bottom of a report is far less informative than the reasoning above it. That is one of the most practical lessons in this book: editors can combine and weigh reasoning, but they cannot do much with a bare verdict. Reviewers also miss errors. In a study published in JAMA in 1998, Fiona Godlee, Catharine Gale, and Christopher Martyn took a paper that had already been accepted by the BMJ, introduced eight deliberate weaknesses in design, analysis, and interpretation, and sent it to hundreds of reviewers. On average, reviewers identified only about two of the eight. A later study led by Sara Schroter, published in the Journal of the Royal Society of Medicine in 2008, inserted nine major errors into three papers and sent them to more than six hundred BMJ reviewers; the average reviewer detected fewer than three of the major errors, and training produced only modest improvement. In emergency medicine, William Baxt and colleagues reported in 1998 on a fictitious manuscript seeded with errors and sent to all reviewers of the Annals of Emergency Medicine; many recommended acceptance despite fundamental flaws. These studies share a design that flatters no one: a manuscript deliberately constructed to be flawed, sent to reviewers who had no reason to expect it. They probably overstate how often errors survive in the ordinary course of review, where several reviewers and an editor scrutinise the work, and where authors are usually not trying to hide anything. But they establish beyond reasonable doubt that individual reviewers, working under normal conditions, miss a large share of serious problems. Systematic reviews of whether editorial peer review improves the quality of published research have reached cautious conclusions. A Cochrane review led by Tom Jefferson in 2007 found little rigorous evidence either way, largely because few studies had been designed to answer the question well. Later work has found that reporting quality tends to improve between submission and publication, but the size of the effect varies and the mechanism is not always clear. What review cannot do The limitations of peer review follow from its structure, and they are worth stating plainly, because a reviewer who expects the process to do something it cannot will review badly. Review cannot detect fraud reliably. A reviewer sees a manuscript, not a laboratory. If an author fabricates data competently, presents internally consistent results, and describes plausible methods, there is usually no way for a reviewer to know. The major fraud cases of the past quarter-century, from the fabricated organic transistor results of Jan Hendrik Schön at Bell Labs, exposed in 2002, to the fabricated stem-cell lines of Hwang Woo-suk, whose papers in Science were retracted in 2006, passed through review at the most selective journals in the world. They were uncovered by readers, colleagues, and whistleblowers after publication, often because someone noticed duplicated figures or impossible consistency across supposedly independent experiments. Chapter 7 returns to what a reviewer can do when something looks wrong. The honest starting point is that review is not designed as a fraud detector and should not be judged as one. Review cannot replicate. The reviewer's evidence is the manuscript and whatever supplementary material the authors provide. Even where data and code are available, reviewers rarely have the time to rerun analyses, and almost never have the resources to repeat experiments. Review assesses whether a claim is plausibly supported by the evidence presented, not whether the claim will hold up. Review cannot compensate for a poor question. A reviewer can point out that a study is underpowered, that a comparison group is inappropriate, or that a conclusion overreaches. They cannot make an uninteresting study interesting, and a report that tries to redesign the whole project is usually of little use to anyone. Review cannot be neutral. Reviewers are human, drawn from the same communities as the authors, with the same rivalries, loyalties, theoretical commitments, and unconscious associations. Chapter 6 examines the evidence on bias in detail. The point here is that a system built on human judgement inherits human failings, and the question is not whether bias exists but how it can be reduced and made visible. Who relies on the verdict The weight carried by the phrase "peer-reviewed" extends far beyond the journal. Systematic reviewers and meta-analysts often restrict their searches to peer-reviewed literature, so a study that passes review becomes a data point in a pooled estimate that may inform a clinical guideline or a regulatory decision. Science journalists use publication in a reviewed journal as a threshold for coverage. Hiring, promotion, and funding committees count reviewed publications, and frequently weigh them by the selectivity of the venue. The law relies on it too. In Daubert v. Merrell Dow Pharmaceuticals, decided by the United States Supreme Court in 1993, the Court set out factors that federal judges may consider when deciding whether expert scientific testimony is admissible. Whether a theory or technique has been subjected to peer review and publication is one of them. The Court was careful to say that publication is not a requirement and that review is an imperfect filter, but the decision nonetheless made peer review a formal consideration in legal proceedings affecting product liability, criminal forensics, and environmental regulation. The COVID-19 pandemic showed both how much the certification function matters and how easily it can be bypassed or undermined. Preprints, which appear without review, became a main channel of scientific communication during 2020, and many were covered by the press before any expert had examined them. At the same time, peer review itself failed conspicuously in at least one high-profile case. In June 2020 The Lancet and the New England Journal of Medicine retracted papers based on a hospital database supplied by the company Surgisphere, after independent researchers raised questions that the company could not answer about where its data had come from. The papers had passed review at two of the most selective medical journals in the world. The data behind them could not be verified. For a reviewer, the lesson of this chain of reliance is not that every report carries the fate of public health. It is that the verdict a reviewer contributes to does not stay inside the journal. Downstream users rarely read the reviews; they see only the outcome. A reviewer who waves a paper through with a careless "minor revision" has, in effect, lent their expertise to every later use of that paper, without the users having any way to know how much scrutiny was really applied. That is one reason, discussed in Chapter 8, why the movement to publish reviews alongside papers has gathered strength: it lets downstream readers see what the certification consisted of. Why it is still worth doing well Given all this, it would be easy to conclude that peer review is theatre, and some critics have said as much. The conclusion does not follow. The evidence shows that review is noisy and fallible, not that it is useless. Manuscripts are routinely improved by it; errors are routinely caught; overstated claims are routinely moderated. Most working scientists can point to a review that saved them from publishing something wrong. And the alternatives that have been proposed, from publishing everything and letting readers sort it out to relying on metrics of attention, have failure modes of their own. The more realistic reform agenda, explored in Chapter 8, keeps expert review but changes when it happens, who can read it, and how reviewers are credited. What the evidence does suggest is that the quality of peer review depends heavily on the quality of individual reviews, and that individual reviews vary enormously. A careful, specific, well-reasoned report can steer a paper and an editor toward a sound outcome. A careless or hostile one can do real damage, delaying good work, letting weak work through, or discouraging a young author from submitting again. The difference between the two lies largely in the habits of the person writing, and those habits can be learned. That is where this book begins its practical work. Before turning to how to read and write a review, however, a reviewer needs to understand the system their report enters: who reads it, who decides, and what they need. That is the subject of the next chapter. A working definition It helps to end with a definition that reflects what the evidence and history suggest, rather than what the ceremony of "peer-reviewed" implies. Peer review is a structured request for expert advice, made by an editor to a small number of specialists, about whether a piece of work is sound, whether it suits a particular venue, and how it might be improved, delivered on the basis of the evidence the authors have chosen to present, under time pressure, and with the knowledge that the final decision belongs to someone else. Each clause of that definition carries an obligation. Structured means the reviewer should organise their advice so it can be used. Expert means the reviewer should confine confident judgements to what they actually know. Advice means the reviewer is not the decision-maker. The evidence the authors have chosen to present means the reviewer should notice what is missing as well as what is there. Under time pressure means the reviewer should prioritise, spending their limited attention where it matters most. The chapters that follow are, in a sense, an extended commentary on these obligations. Chapter 2. The Editorial System: Who Decides What A reviewer's report is one input into a decision made by someone else, and it travels through a system that most reviewers never see. Understanding that system is not administrative trivia. It determines what a report should contain, how its recommendation will be read, which remarks the authors will see, and what happens when a reviewer raises a concern about something other than the science. This chapter follows a manuscript from submission to decision, describing the people it passes through and what each needs. The journey of a manuscript A manuscript arrives through an online submission system, usually one of a small number of commercial platforms that most journals license. Before any scientist looks at it, editorial staff typically perform technical checks: that the required files are present, that the word count and formatting are within limits, that ethics approvals and data availability statements have been provided, that author details and conflict of interest declarations are complete. Many publishers now also run automated screening at this stage, including similarity checks for text overlap with published work and, increasingly, tools designed to detect manipulated images or patterns associated with commercially produced fraudulent papers. The manuscript then reaches an editor. At large journals this may be a professional editor, a full-time employee with a doctorate who no longer runs a laboratory; at most society and specialist journals it is an academic editor, a working scientist who handles manuscripts alongside research and teaching. The first decision this editor makes is whether to send the paper for review at all. Desk rejection, the return of a manuscript without external review, is common at selective journals, where a large majority of submissions may be declined at this stage. The reasons are usually scope, perceived significance, or obvious problems of quality or presentation. Desk rejection spares authors weeks of waiting and spares reviewers from assessing work the journal would never publish, which is why, uncomfortable as it is, it is generally considered good editorial practice when done promptly and with a brief explanation. If the paper goes forward, an editor identifies reviewers. At many journals an editor-in-chief assigns the manuscript to a handling editor, sometimes called an associate or academic editor, who has relevant expertise and manages the rest of the process. The handling editor searches for reviewers using their own networks, the manuscript's reference list, databases of past reviewers, and, at some journals, algorithmic reviewer-suggestion tools. Authors are often invited to suggest or exclude reviewers. Editors vary in how they treat these suggestions; some avoid author-suggested reviewers entirely, particularly since investigations in the mid-2010s revealed peer-review rings in which authors supplied contact details for fake reviewer accounts they controlled, leading to the retraction of scores of papers across several publishers. Finding reviewers is frequently the slowest step. Editors routinely send many invitations for each acceptance, and reviewer fatigue, the concentration of reviewing burden on a relatively small pool of willing and reliable people, is a widely recognised problem. This is one reason editors are often glad to invite early-career researchers, whose expertise in current techniques is frequently more precise than that of senior investigators and whose availability is sometimes greater. Once reviews are in, the handling editor reads them alongside the manuscript and reaches a decision, or makes a recommendation to the editor-in-chief who makes it. The decision letter goes to the authors with the reviewers' comments attached. If revisions are requested, the revised manuscript may return to the same reviewers, to a subset of them, or only to the editor, depending on the extent of the changes and the journal's policy. The roles and their obligations Each position in this system carries distinct responsibilities, and several of the most common problems in peer review arise when someone acts outside their role: a reviewer who treats their report as a final verdict, an editor who forwards reports without reading them, an author who treats review as an adversarial negotiation. The main roles and their core duties are set out in Table 1. Table 1. Principal roles in journal peer review and their core obligations. Role Primary responsibility Key ethical duties Typical failure Author Report the work fully and accurately Honest data, disclosure of conflicts, credit to contributors Overclaiming; withholding data Handling editor Manage review and reach or recommend a decision Choose qualified, unconflicted reviewers; weigh reports critically Treating reviews as votes Editor-in-chief Set policy and take final responsibility Consistency, appeals, handling misconduct Favouring prominent authors Reviewer Advise on soundness, significance, and improvement Confidentiality, disclosure, fairness, candour about expertise Harshness; scope creep Publisher and staff Operate systems and integrity checks Screening, record-keeping, corrections Opaque processes The table compresses a good deal, and two rows deserve fuller comment. The handling editor is the person a reviewer is really writing for. They are usually an expert in the broad field, but not necessarily in the specific technique or subfield of the manuscript; that is why they sought reviewers. They will read two or three reports, often of very different lengths, tones, and recommendations, and must reconcile them. What helps them most is a report that makes its reasoning visible, distinguishes the problems that would change the conclusions from those that would merely improve the presentation, and is candid about the reviewer's own confidence. What helps them least is a verdict without reasons, or a list of forty objections of undifferentiated weight. Good editors do not treat reviews as votes. If two reviewers recommend minor revision and a third identifies a fundamental flaw that the others missed, a careful editor will weigh the substance of the third report, not count the recommendations. Equally, if one reviewer recommends rejection on grounds that are clearly a matter of taste or theoretical allegiance, the editor may discount that recommendation. This is why the reasoning in a review matters more than the recommendation. Reviewers sometimes feel frustrated when an editor's decision does not follow their advice, but the system is designed that way: the reviewer advises on the evidence they have, and the editor decides with the benefit of all the evidence. The editor-in-chief holds ultimate responsibility for what the journal publishes and for its policies, and is usually the person who handles appeals and serious concerns about misconduct. Reviewers rarely deal with the editor-in-chief directly, but it is useful to know that the handling editor is not the last line of escalation. If a reviewer's serious concern is ignored, the editor-in-chief, and beyond them the publisher's research integrity team, are the appropriate next contacts. What the editor's invitation is asking A review invitation is a specific request, and reading it carefully is the first act of good reviewing. Most invitations include the title and abstract, a deadline, and sometimes a note from the editor about what they would like the reviewer to focus on. That note matters. An editor who writes "we would particularly value your assessment of the single-cell analysis" is telling the reviewer that someone else will cover the physiology. A reviewer who ignores the request and writes a general report may duplicate another reviewer's work while leaving the question the editor actually needed answered. Three questions should be answered honestly before accepting. First, is this within my competence? A reviewer does not need to be expert in every aspect of a manuscript, but they should be able to assess a substantial part of it with confidence, and they should be willing to say which parts they cannot assess. Accepting a review for which one is not qualified and then writing confidently about it is a quiet form of misrepresentation. Second, do I have a conflict of interest? Chapter 5 treats this in detail. At this stage, the question is whether any relationship with the authors, the work, or its subject would lead a reasonable observer to doubt the reviewer's impartiality. If so, the reviewer should either decline or disclose the relationship to the editor and let them decide. Third, can I meet the deadline? Late reviews are one of the most common complaints from both editors and authors, and they cost authors real time at stages of their careers when time matters greatly. It is better to decline promptly, perhaps suggesting a qualified colleague, than to accept and then delay. If circumstances change after accepting, a short message to the editor asking for an extension is courteous and usually granted. Early-career researchers sometimes receive review requests indirectly, when a senior colleague asks them to help with a review the senior colleague has been invited to write. This practice, sometimes called ghostwriting of reviews, is widespread. A survey led by Gary McDowell and colleagues, published in eLife in 2019, found that many early-career researchers had co-reviewed with a principal investigator, and that a large share of those had not been named to the editor. The authors argued that this deprives junior researchers of credit and misleads editors about who actually assessed the work. Most journals permit co-reviewing, but they expect the invited reviewer to obtain permission before sharing a confidential manuscript and to name the co-reviewer. An early-career researcher asked to help should ask whether the editor has been told; a principal investigator who asks for help should tell the editor and ensure the junior colleague is credited. What happens to a report Once submitted, a report is usually divided into two parts: comments for the authors and confidential comments for the editor. The authors will see the first part verbatim, sometimes lightly edited by the editor to remove anything inappropriate. They will not see the confidential comments. The division is useful but frequently misused. Confidential comments are appropriate for information the editor needs but the authors should not see: a note about the reviewer's own limits of expertise, a concern about possible duplicate publication or data manipulation that needs investigation before the authors are approached, a candid view of whether the work reaches the journal's bar for significance, or an explanation that the reviewer has a relationship with the authors that they judged not to be disqualifying. They are not appropriate for delivering criticisms that the reviewer is unwilling to make openly. A reviewer who writes a mild report to the authors and a damning note to the editor leaves the authors unable to respond to the real reasons for rejection. The general principle is that any substantive criticism of the science should appear in the comments to authors, where it can be answered. Most journals send each reviewer the other reviewers' reports and the decision letter after a decision is made. Reading these is one of the best ways for a new reviewer to calibrate. It shows what others noticed that one missed, how different reviewers weighted the same problems, and how the editor combined the advice. Where a reviewer's view differed sharply from the decision, it is worth considering whether the editor saw something in the other reports that justified the difference. If the manuscript is revised, the reviewer may be asked to assess the revision. The task then changes. The question is whether the authors have adequately responded to the points raised, not whether the reviewer can find new objections. Raising entirely new issues at the second round, when they could have been raised at the first, is a common source of author frustration and prolongs review without clear benefit. New problems introduced by the revision itself are fair game, and so are serious problems that were genuinely missed the first time, but the reviewer should acknowledge that they are new. Decisions and their meaning Journals use a small set of decision categories, though the labels vary. Accept is rare after first review. Minor revision means the paper is fundamentally sound and needs only limited changes, often without further external review. Major revision means substantial problems exist but appear fixable; the revised paper will usually be re-reviewed. Reject and resubmit, used by some journals, means the problems are serious enough that a substantially new manuscript is required, though the journal is willing to consider it. Reject means the journal will not consider the work further, whether because of fundamental flaws or insufficient fit or significance. A reviewer's recommendation should correspond to these meanings, and the report should make clear which category the reviewer has in mind and why. It is particularly helpful to state what would be required to change the recommendation. "I would support publication if the authors can show that the effect survives correction for multiple comparisons and holds in the second cohort" gives an editor far more to work with than "major revision". Reviewers should also remember that selection criteria vary between journals and that their recommendation is specific to the venue. A paper that is too incremental for a broad-scope journal may be an excellent contribution to a specialist one. A good report makes this explicit, separating the reviewer's view of the work's soundness, which should not change from journal to journal, from their view of its fit, which should. Variations on the standard workflow Not every manuscript follows the path just described, and a reviewer should recognise the variants, because each changes what the report is for. Registered Reports move the main round of review before the results exist. Introduced at the journal Cortex in 2013, under the editorial leadership of the psychologist Chris Chambers, and now offered by hundreds of journals, the format asks authors to submit their introduction, hypotheses, methods, and analysis plan before collecting data. Reviewers assess this Stage 1 protocol on the importance of the question and the rigour of the design. If the protocol is accepted in principle, the journal commits to publishing the final paper regardless of whether the results support the hypotheses, provided the authors follow the approved plan. At Stage 2, reviewers check that the plan was followed and that the conclusions follow from the results. For a reviewer, the shift is substantial. At Stage 1 there is no result to be impressed or disappointed by, so the assessment rests entirely on whether the design can answer the question. At Stage 2 the reviewer's freedom is deliberately restricted: they are not invited to reject the paper because the findings are unexciting or unexpected. The format was designed to counter publication bias, and it works only if reviewers respect its rules. Cascading or transfer review allows a manuscript rejected by one journal to be passed, with its reviews, to another journal from the same publisher, or in some arrangements to a different publisher altogether. Reviewers are often asked at the time of review whether they consent to their report being transferred. Consent costs little and can save authors months, but it means a report should be written so that it remains useful to an editor at a different venue; a report whose substance is simply "not important enough for this journal" transfers poorly. Special issues and conference proceedings often run on compressed timelines, with guest editors who may be less experienced than a journal's regular editorial board. In recent years several publishers have retracted large numbers of papers from special issues after discovering that guest editors had been impersonated or that review had been compromised. A reviewer invited to a special issue should apply exactly the same standards as for a regular submission and should be alert, as any reviewer should, to signs that a manuscript has been produced by a paper mill, a topic returned to in Chapter 7. Appeals and disputes Authors who believe a decision was wrong may appeal. The Committee on Publication Ethics, known as COPE, an organisation founded in 1997 whose guidance most major publishers follow, recommends that journals have a clear appeals process. Appeals are usually handled by the editor-in-chief or a senior editor, who may seek a further opinion. Reviewers are occasionally asked to respond to an author's rebuttal. When that happens, the right approach is to read the rebuttal on its merits, concede points where the authors are right, and explain clearly where the original concern still stands. A reviewer who treats an appeal as a personal challenge has forgotten that their role was always advisory. The editorial system, then, is designed to distribute responsibility. Authors are responsible for the honesty and completeness of what they submit. Reviewers are responsible for accurate, candid advice within their competence. Editors are responsible for choosing reviewers well, weighing their advice critically, and making decisions they can defend. When each role is performed well, the system can reach good decisions despite the noise in any single review. When a reviewer understands that their report is one part of this structure, they can write it to be used, which is the subject of the next two chapters. Chapter 3. Reading a Manuscript Critically Most weak reviews fail before a word is written. The reviewer reads the manuscript once, from beginning to end, marking objections as they arise, and then assembles those objections into a report. The result is predictable: a long list of comments of wildly varying importance, heavily weighted toward the introduction and early methods where the reviewer's attention was freshest, often missing the one problem in the analysis that actually determines whether the conclusions hold. Linear reading is how people read papers for pleasure or for background. It is not how to assess one. A more reliable approach treats reading as a sequence of distinct passes, each with its own purpose. The first establishes what the paper claims. The second tests whether the design and analysis can support those claims. The third examines the details: reporting, figures, statistics, data, and the fit between what was done and what is said. The passes need not be rigid, and experienced reviewers blend them, but separating them at first builds the habit of asking the important questions before the easy ones. The first pass: what is being claimed? Before assessing whether a paper is right, a reviewer must know precisely what it says. This sounds trivial and is not. Many manuscripts contain several claims of different strength, and the abstract, the discussion, and the title frequently state them differently. A title may announce that a gene "controls" a behaviour while the results show a correlation in one mouse strain; an abstract may report a "significant reduction" while the discussion concedes that the effect was small and confined to a subgroup. The first pass should therefore be quick and focused on the architecture of the argument. Read the title, abstract, and final paragraphs of the introduction, where the authors usually state their aims. Look at every figure and table, with their legends, without reading the results text. Then read the discussion's opening and closing paragraphs. At the end of this pass, a reviewer should be able to write down, in two or three sentences, the central claim of the paper, the key evidence offered for it, and the type of study that produced that evidence. It is worth actually writing this summary. It becomes the opening of the review, where it serves an important function discussed in the next chapter: it shows the authors and the editor that the reviewer has understood the work, and it exposes any misunderstanding before it contaminates the rest of the assessment. It also serves the reviewer. If the central claim cannot be stated clearly after a careful first pass, that is itself a finding: the manuscript is unclear about what it is arguing, and saying so is one of the most useful things a report can do. The first pass should also identify the claim's type. Is it causal, or descriptive, or predictive? Is it a claim about a mechanism, a population, a method's performance, or the existence of a phenomenon? The type determines what evidence is required. A causal claim from observational data needs a credible strategy for addressing confounding; a claim about a new method's superiority needs a fair comparison against current alternatives on relevant benchmarks; a claim about a population needs a sample that represents it. Much of the substance of a good review consists of matching the claim to the evidence its type requires. The second pass: can the design support the claim? With the claim in hand, the reviewer reads the methods and results closely, asking a single overarching question: if everything reported here was done exactly as described, would it justify the conclusion? This question sets aside, for now, whether things were done as described. It concerns the logic of the study. Several recurring problems are worth checking for explicitly, because they are common and because they are easy to miss when reading in the authors' frame. Alternative explanations. For every main result, ask what else could have produced it. In experimental work this usually means asking whether the controls rule out the obvious alternatives: whether a knockdown effect might be off-target, whether a drug effect might be due to the vehicle, whether a behavioural difference might reflect a motor deficit rather than a cognitive one. In observational work it means asking about confounding, selection, and reverse causation. A reviewer who can name a specific, plausible alternative explanation that the design does not exclude has found something important. A reviewer who says only that "other factors may be involved" has not. Comparison and baseline. Every claim of an effect is a claim about a difference, and the choice of comparison determines what the difference means. Is the control group appropriate? Is a new method compared with the current best alternative or with a straw man? Is a change measured against a baseline that could itself have shifted? Sample and power. Is the sample large enough to detect effects of the size claimed, or to make a null result informative? Small studies that report large effects deserve particular scrutiny, because they are disproportionately likely to be false positives or inflated estimates, a point made forcefully by John Ioannidis and by Katherine Button and colleagues, whose 2013 analysis in Nature Reviews Neuroscience found the median statistical power of neuroscience studies to be low. Where sample sizes are not justified by a power calculation or equivalent reasoning, a reviewer may reasonably ask for one. Independence of observations. A frequent and consequential error is to treat non-independent measurements as independent: multiple cells from the same animal, repeated measures from the same participant, several samples from the same culture. Doing so inflates apparent sample sizes and produces spuriously small p-values. It is worth checking what the unit of analysis is and whether it matches the unit of replication. Analytical flexibility. Were the analyses planned in advance, or chosen after seeing the data? Were outcomes, covariates, or exclusion criteria selected in ways that might have favoured the reported result? In clinical trials, comparison with the registered protocol is often revealing; in other fields, preregistration is increasingly common and should be checked when it exists. Where there is no preregistration, the reviewer can still ask whether the results are robust to reasonable alternative analytical choices. Generalisation. Do the conclusions extend beyond the conditions actually tested? Findings in one cell line, one species, one population, or one dataset are frequently described as if they held generally. A reviewer should ask the authors to match the scope of their claims to the scope of their evidence. By the end of the second pass, a reviewer should know whether there are any problems that, if not addressed, would mean the central conclusion is unsupported. These are the major concerns. They may be few, and there may be none. Identifying them correctly is the most valuable thing a reviewer does. The third pass: details, reporting, and data The third pass is where most reviewers spend most of their time, which is why it belongs last. Here the reviewer checks whether things were done and reported as they should be: whether statistical tests are appropriate and correctly described, whether figures display what the text says, whether the numbers in different places agree, whether the methods are described in enough detail to be reproduced, and whether data and code are available as the journal requires. Reporting guidelines are a practical aid here. Over the past three decades, groups of methodologists and editors have produced checklists specifying what should be reported for common study designs, and many journals require authors to complete the relevant checklist on submission. A reviewer does not need to audit every item, but knowing the relevant guideline helps identify omissions quickly. The most widely used are summarised in Table 2; the EQUATOR Network, an international initiative founded in 2008 to improve the reliability of health research reporting, maintains a searchable library of several hundred more. Table 2. Widely used reporting guidelines and the study types they cover. Guideline Study type Issued by or associated with Useful for checking CONSORT Randomised controlled trials CONSORT Group Randomisation, allocation, flow of participants PRISMA Systematic reviews and meta-analyses PRISMA Group Search strategy, selection, risk of bias STROBE Observational epidemiology STROBE Initiative Confounding, selection, missing data ARRIVE Animal research NC3Rs (UK) Randomisation, blinding, sample size STARD Diagnostic accuracy studies STARD Group Reference standard, spectrum of patients The guidelines matter to review for a reason beyond completeness. Their items encode the specific ways each type of study tends to go wrong. ARRIVE asks whether animals were randomly allocated and whether outcome assessors were blinded because unblinded, unrandomised animal studies have repeatedly been shown to report larger effects. CONSORT asks for a participant flow diagram because attrition that differs between arms can create an apparent treatment effect. Reading a manuscript against the relevant guideline is, in effect, reading it with the accumulated experience of the methodologists who studied that design's failures. Figures deserve particular attention in the third pass. Check that axes are labelled and scaled honestly, that error bars are defined, that the sample size for each panel is stated, and that representative images are accompanied by quantification. Look for panels that seem too clean or that resemble each other in ways they should not; Chapter 7 discusses what to do if something appears duplicated or altered. Check that individual data points are shown where sample sizes are small, since bar charts of means can conceal very different distributions. Numbers should be checked for internal consistency. Do the sample sizes in the methods match those in the figure legends and tables? Do percentages add up? Are reported test statistics, degrees of freedom, and p-values consistent with one another? Tools exist to help with the last question. The program statcheck, developed by Michèle Nuijten and colleagues, recomputes p-values from reported test statistics in APA-formatted text, and the GRIM test, proposed by Nick Brown and James Heathers in 2016, checks whether reported means are arithmetically possible given the sample size and the granularity of the underlying data. Inconsistencies found this way are usually typographical, but they can indicate deeper problems, and in either case they should be corrected. Finally, the reviewer should check what the data and code availability statement says, and whether it is true. A statement that data are "available on reasonable request" is weaker than a deposit in a public repository, and studies have found that such requests are frequently unanswered. Many funders and journals now require deposit where ethically possible. If the journal's policy requires data to be available and it is not, or if a link leads nowhere, the reviewer should say so. If the data are available and the reviewer has the time and competence, looking at them, even briefly, can reveal problems invisible in the manuscript. The three passes in practice A hypothetical example shows how the passes change what a reviewer notices. Suppose a manuscript reports that a dietary supplement improves memory in older adults. Its title says the supplement "enhances cognitive function in ageing". The abstract reports a randomised trial of 120 participants over twelve weeks and a significant improvement on a word-recall task. A linear reader might begin by objecting to the introduction's selective citations, move on to request more detail about the supplement's formulation, question the choice of font in a figure, and eventually, somewhere in the results, notice that the primary outcome listed in the trial registry was a composite cognitive score, not word recall. That last observation, buried as comment seventeen of twenty-two, is the one that matters. The first pass, done properly, would have surfaced it early. Writing down the central claim forces the question of which outcome supports it, and the type of claim, a causal effect from a randomised trial, directs the reviewer immediately to CONSORT's concerns: was the primary outcome prespecified, and was it the one reported? The second pass would then ask whether word recall was among several secondary outcomes, whether the composite score showed any effect, whether multiple comparisons were accounted for, and whether a twelve-week improvement on one task justifies a title claiming enhanced "cognitive function" in general. It would also ask about blinding, since a supplement with a distinctive taste might allow participants to guess their allocation, and about attrition, since older participants who experience side effects may drop out unevenly between arms. Only in the third pass would the reviewer check the formulation details, the consistency of the numbers across tables, and the figure design. Those checks are still worth doing. But the report built from these passes leads with the outcome-switching question and the overreach in the title, which are the problems that determine whether the conclusion stands, and relegates the rest to minor comments. The editor reading it knows at once what the decision turns on. The authors know what they must address, and they may well have a good answer: perhaps the registry was amended before unblinding, for documented reasons. A review that asks the right question clearly allows that answer to come out. Reviewing within one's competence Few reviewers are expert in every method a modern manuscript uses. A single paper in cell biology may combine imaging, genomics, animal behaviour, and sophisticated statistics; a paper in ecology may combine field sampling, remote sensing, and Bayesian modelling. The reviewer's obligation is not to be omniscient but to be clear about where their competence ends. The practical rule is to review confidently what one knows, to review cautiously what one partly knows, and to state plainly what one cannot assess. "I am not able to evaluate the phylogenetic methods and would recommend that the editor seek a specialist opinion on them" is a useful sentence. It tells the editor precisely where a gap exists in the assessment, and editors frequently act on it. By contrast, a reviewer who comments confidently on methods they do not understand may mislead the editor, burden the authors with inappropriate requests, or miss real problems while inventing false ones. Statistics deserve a special word. Many reviewers who are expert in their experimental domain are less secure in statistical analysis, and statistical problems are among the most common serious flaws in published research. A reviewer who is uncertain whether an analysis is appropriate should say so and suggest statistical review. Several journals employ statistical reviewers or editors for exactly this reason, and an editor who is told that a statistical question needs specialist attention can obtain it. Time, attention, and proportion A thorough review of a substantial manuscript takes time, typically several hours spread over more than one sitting. There is no merit in taking longer than necessary, but there is real cost in taking less than the work requires. The distribution of time matters as much as its total. A useful rough allocation is to spend a modest share on the first pass, the largest share on the second, and the remainder on the third, adjusting for the paper's complexity. Reviewers who spend most of their time correcting typographical errors and requesting additional citations have allocated their attention poorly, however long they spent. It is also worth leaving time between reading and writing. First reactions to a manuscript are often stronger than considered ones, in both directions. A day's distance frequently turns an objection that seemed decisive into one that is real but minor, or reveals that an apparently convincing result depends on an assumption not stated. The reviewer who writes immediately after a single read is likely to produce a report that reflects their mood as much as the manuscript. Reading with charity The final principle of critical reading is one that sounds soft but is methodological: read the manuscript as the strongest version of what the authors are trying to say. When a passage is ambiguous, consider the most reasonable interpretation before assuming the worst one. When a control seems missing, check whether it appears in the supplementary material or whether a different control addresses the same concern. When a result seems implausible, ask what would have to be true for it to be correct. Charitable reading is not credulity. Its purpose is to ensure that criticisms, when they come, are aimed at what the authors actually did rather than at a misreading. A report that attacks a straw man wastes the authors' time, undermines the reviewer's credibility with the editor, and often leaves the real weaknesses untouched. A report that engages with the strongest version of the argument, and still finds problems, is one that the authors must take seriously. With the reading done, the reviewer knows what the paper claims, whether the design can support it, and where the details fall short. The next task is to turn that knowledge into a report that an editor can use and authors can act on. Hashtags: #ThePeerReviewEcosystem #PeerReview #ScholarlyPeerReview #EditorialEthics #ResearchIntegrity #ConstructiveCritique #EditorialRoles #ReviewerResponsibilities #HandlingEditors #EditorsInChief #ReviewerEthics #ConflictOfInterest #Confidentiality #ReviewerBias #ImplicitBias #PublicationEthics #COPE #CriticalManuscriptReview #ResearchQuality #ScientificPublishing #OpenPeerReview #RegisteredReports #EditorialDecisionMaking #ResearchTransparency #FutureOfPeerReview
- Functional Magnetic Resonance Imaging (Experimental Paradigm Design and BOLD Analysis)
Download the Book (PDF): Introduction In the autumn of 2009, a poster at a neuroimaging conference reported that a dead Atlantic salmon, placed in a scanner and "shown" photographs of people in social situations, displayed task-related activity in its brain cavity. The authors, Craig Bennett and colleagues, had not discovered piscine social cognition. They had run a standard analysis on a real dataset without correcting for the tens of thousands of statistical tests that a whole-brain map entails, and the noise had obligingly arranged itself into a blob. The poster was a joke with a serious point, and it became one of the most cited cautionary tales in the field. Seven years later a less humorous paper made a related point at scale. Anders Eklund, Thomas Nichols and Hans Knutsson took resting-state scans from hundreds of healthy people, invented fake task timings, and ran them through the most widely used software packages. Under some common settings, the proportion of analyses that reported at least one significant cluster where none could exist was far above the nominal five percent. Four years after that, seventy independent teams were given the same fMRI dataset and the same nine hypotheses; they reached materially different conclusions on several of them, because each team had made a different set of defensible analytic choices. None of these episodes shows that functional magnetic resonance imaging does not work. Each shows something narrower and more useful: that the distance between a person lying in a scanner and a coloured map in a journal is long, that every step along it involves a decision, and that the decisions interact. A map is an inference, not a photograph. Whether the inference is sound depends on how the experiment was designed, how the slow vascular signal was modelled, how motion and physiology were handled, and how the statistics were carried from individual brains to a population claim. This book is about that chain. Its controlling argument is simple to state and easy to forget: in fMRI, the validity of a finding is fixed mainly at the design stage, and analysis can only preserve, never manufacture, the information the design built in. The blood-oxygen-level-dependent (BOLD) signal is sluggish, indirect, small relative to noise, and entangled with head motion, breathing and heartbeat. An experiment that does not anticipate those properties produces data in which the question of interest is confounded with something else, and no amount of downstream sophistication can separate what the design fused. An experiment that does anticipate them produces data in which even a modest analysis can give a trustworthy answer. The practical upshot is that the most consequential decisions in an fMRI study are made before the first participant is scanned: the choice of block or event-related structure, the timing and ordering of trials, the contrasts that will be tested, the tolerance for motion, the sample size, and the inferential procedure that will be applied to the final map. What the reader will find The first two chapters establish what is being measured. Chapter 1 explains the physics and physiology of the BOLD signal: why deoxygenated haemoglobin makes the magnetic resonance signal decay faster, why a burst of neural activity paradoxically produces an excess of oxygenated blood, and what this means for the spatial and temporal precision of anything fMRI can report. Chapter 2 treats the hemodynamic response function, the characteristic rise, peak and undershoot that follows a brief neural event, and the assumptions of linearity and time-invariance on which almost every analysis depends. It also confronts the awkward fact that the response differs across brain regions, people and populations, and considers how much that matters. Chapters 3 and 4 are about design. Block designs, which alternate sustained periods of one condition with another, remain the most statistically powerful way to detect whether a region responds at all. Event-related designs, which present brief trials in an interleaved and often randomised order, sacrifice some of that power to gain the ability to separate trial types, sort trials by behaviour after the fact, and estimate the shape of the response. Mixed designs try to have both. The chapters work through the concept of design efficiency, the trade-off between detecting an effect and estimating its time course, the role of jittered intervals, and the logic of subtraction and its alternatives. Readers who plan experiments will find the most direct practical guidance here. Chapter 5 turns to the single-subject general linear model, the workhorse of task fMRI. It explains how a design matrix is built by convolving stimulus timings with a response model, why deconvolution is possible at all, how contrasts are specified, and why the temporal autocorrelation of fMRI noise has to be modelled rather than ignored. It also covers the pitfalls of correlated regressors, parametric modulation and trial-wise estimation. Chapter 6 deals with the problem that dominates practical fMRI more than any other: head motion, and its cousins, cardiac and respiratory noise. It explains why a movement of a fraction of a millimetre can produce apparent "activation", why motion that is correlated with the task is so much more dangerous than motion that is not, and how realignment, nuisance regression, censoring, component-based denoising, multi-echo acquisition and, above all, prevention fit together. Chapters 7 and 8 carry the analysis to the group. Chapter 7 covers spatial normalisation, smoothing and the multilevel logic by which individual estimates become population inferences, including the reasons fixed-effects analyses cannot generalise and the uncomfortable arithmetic of statistical power in typical fMRI samples. Chapter 8 addresses statistical parametric mapping proper: the multiple-comparisons problem, random field theory, cluster-level inference and its failures, false discovery rate, permutation methods, and the reporting and sharing practices that now distinguish credible work from the rest. The conclusion does not summarise. It argues that the lessons of the preceding chapters add up to a particular way of planning a study, one that treats the analysis as something to be designed alongside the paradigm rather than chosen afterwards, and it identifies what remains genuinely unsettled. Who the book is for The intended reader is someone who needs to design, run, analyse or critically read an fMRI study: a graduate student beginning a first project, a researcher from psychology, medicine or computer science moving into imaging, a clinician evaluating a report, or a reviewer who wants to know which questions to ask of a manuscript. No prior knowledge of magnetic resonance physics is assumed, though comfort with basic statistics, especially regression and hypothesis testing, will help. Mathematical ideas are explained in words and with a minimum of notation. Where a formula is unavoidable it is stated in plain terms. The book deliberately does not cover every use of fMRI. Resting-state functional connectivity, multivariate pattern analysis, representational similarity analysis, effective connectivity and real-time neurofeedback each deserve their own treatment. They appear only where they bear directly on the core problems of task design and BOLD analysis, which remain the foundation on which those other methods are built. Nor is this a software manual. The major packages, SPM, FSL and AFNI, together with newer pipelines such as fMRIPrep, are mentioned where their defaults or conventions matter to an argument, but the principles apply regardless of which tool a laboratory uses. A note on evidence fMRI methodology has a large and argumentative literature, and some of its best-known debates are still live. Where the book reports a specific finding, it names the study. Where consensus exists, it says so; where it does not, it lays out the competing positions and what would settle them. A short list of notes and a guide to further reading appear at the end for readers who want to follow any thread back to its source. The field has changed a great deal since the first BOLD images were published in the early 1990s. Scanners are faster, software is more robust, datasets are shared openly, and the community's tolerance for unreplicable results has dropped sharply. The central difficulty has not changed. The scanner measures blood, the question concerns minds, and the bridge between them is built out of design decisions. Building that bridge well is what this book is about. Chapter 1. What the BOLD Signal Measures Every fMRI finding rests on a physical fact discovered long before anyone imagined looking at thought: haemoglobin changes its magnetic character depending on whether it carries oxygen. In 1936 Linus Pauling and Charles Coryell showed that oxygenated haemoglobin is diamagnetic, weakly repelled by a magnetic field, while deoxygenated haemoglobin is paramagnetic, weakly attracted to one. Half a century later, Seiji Ogawa and colleagues at AT&T Bell Laboratories used this property to produce images of rat brains in which blood vessels darkened or brightened depending on the oxygenation of the blood inside them. Their 1990 paper in the Proceedings of the National Academy of Sciences named the effect blood-oxygenation-level-dependent contrast. Within two years, three groups, led by Kenneth Kwong at Massachusetts General Hospital, Ogawa himself with collaborators in Minnesota, and Peter Bandettini at the Medical College of Wisconsin, had published BOLD images of the working human brain responding to visual stimulation and finger movement. No injected tracer was needed. The blood was its own contrast agent. Understanding what that signal is, and what it is not, is the first design decision, because every later choice about timing, resolution and interpretation follows from it. From magnetism to image intensity A magnetic resonance image is built from the behaviour of hydrogen nuclei, mostly in water, placed in a strong static magnetic field. The nuclei align slightly with the field, are tipped away from alignment by a radiofrequency pulse, and then precess back, emitting a signal as they do. How quickly that signal fades depends on the local environment. The decay that matters for fMRI is called T2*, and it is sensitive to tiny inhomogeneities in the magnetic field within each voxel. When nuclei in the same voxel experience slightly different field strengths, they precess at slightly different rates, fall out of phase with one another, and their combined signal cancels faster. Deoxygenated haemoglobin, being paramagnetic, distorts the field around the red blood cells and vessels that contain it. More deoxyhaemoglobin in a voxel means more field distortion, faster dephasing, a shorter T2, and therefore a darker voxel in an image acquired with sensitivity to T2. Less deoxyhaemoglobin means a brighter voxel. The BOLD signal is therefore, in the first instance, a measure of the concentration of deoxygenated haemoglobin in a voxel, and it is an inverse measure: the signal goes up when that concentration goes down. To catch these small differences, fMRI uses gradient-echo echo-planar imaging, a technique proposed by Peter Mansfield in 1977 that reads out an entire two-dimensional slice after a single excitation, in tens of milliseconds. The echo time, the delay between excitation and readout, is chosen to be close to the tissue's T2* so that the image is maximally sensitive to changes in it; at 3 tesla, typical echo times are around thirty milliseconds. A full brain volume is assembled slice by slice, and the time taken to acquire one volume, the repetition time or TR, has historically been about two seconds. Simultaneous multi-slice or "multiband" acquisition, developed in the late 2000s by groups including David Feinberg's and Kamil Uğurbil's, excites several slices at once and now routinely brings whole-brain TRs below one second. The same sensitivity to field inhomogeneity that makes BOLD possible also produces its best-known artefact. Where air meets tissue, as in the sinuses above the orbitofrontal cortex and the ear canals beside the inferior temporal lobes, the magnetic field is badly distorted. The signal in these regions drops out or is displaced. An experimenter interested in reward valuation or semantic memory, both of which involve precisely these areas, needs to know this before designing the study, because a region that cannot be imaged well cannot show an effect, and the absence of an effect there means nothing. Slice orientation, thinner slices, shorter echo times, and field-map-based distortion correction all help; none eliminates the problem. Why active brain tissue looks brighter The paradox at the heart of BOLD is that active neurons consume oxygen, yet active brain regions show less deoxyhaemoglobin, not more. The resolution came from positron emission tomography in the 1980s. Peter Fox and Marcus Raichle, measuring blood flow and oxygen metabolism in human visual and somatosensory cortex, reported in 1986 that stimulation raised regional cerebral blood flow by roughly thirty to fifty percent, while the rate of oxygen consumption rose by only about five percent. The vascular system oversupplies. Fresh, fully oxygenated blood arrives faster than the extra oxygen can be extracted, so the fraction of haemoglobin that is deoxygenated falls, and the T2*-weighted signal rises. Why the brain does this is still debated. One long-standing view held that flow increases in order to meet energy demand, with the overshoot a side effect of the physics of oxygen diffusion from capillaries. A more recent view, set out by David Attwell and Costantino Iadecola among others, holds that blood flow is controlled largely in a feed-forward manner by signalling molecules released during synaptic activity, particularly glutamate-triggered pathways involving nitric oxide, prostaglandins and astrocytes, rather than by a feedback signal reporting energy shortfall. On this account, the hemodynamic response is triggered by the neurotransmission itself, and is only loosely tied to the metabolic cost of that transmission. Work by Catherine Hall and colleagues in 2014 suggested that capillaries, dilated by contractile cells called pericytes, respond earlier than arterioles, adding a further layer to the vascular machinery. For the experimenter, the precise mechanism matters less than three consequences. First, BOLD is a measure of vascular response, not of neural firing, and the relation between the two is mediated by a complex, partly understood chain of cellular events. Second, anything that changes the vasculature or the coupling between neurons and vessels, including age, disease, medication, caffeine, alcohol and carbon dioxide levels, can change the BOLD response without any change in neural activity. Third, the vascular response is slow, which is the subject of the next chapter. Which neural activity does BOLD reflect? The question of what aspect of neural activity drives the hemodynamic response received its most influential answer in 2001, when Nikos Logothetis and colleagues recorded BOLD signals and intracortical electrophysiology simultaneously in the visual cortex of anaesthetised monkeys. They separated the electrical signal into multi-unit activity, which reflects the spiking output of neurons near the electrode, and local field potentials, which reflect the summed synaptic and dendritic currents in a larger volume. The BOLD signal was predicted somewhat better by local field potentials than by spiking; in some recordings, the spiking response adapted and faded while both the field potential and the BOLD response persisted. The standard summary is that BOLD reflects the input to a region and its local processing more than its output. The summary is useful but should not be pushed too hard. Spiking and field potentials are usually correlated, so in most circumstances a region that fires more also shows more BOLD. The dissociation matters in specific situations: where strong input is followed by inhibition, where a region receives modulatory input that changes its sensitivity without changing its firing, or where recurrent local processing is heavy. Inhibition is an especially important case. Inhibitory synaptic activity is metabolically costly and can drive vascular responses, so an increase in BOLD does not necessarily mean that a region is producing more output. It may mean the region is being actively suppressed. The opposite sign is also informative. Sustained negative BOLD responses, in which the signal falls below baseline during stimulation, were characterised in human visual cortex by Amir Shmuel and colleagues in 2002, who found them in regions adjacent to the stimulated part of the visual field. Later work in monkeys linked such negative responses to reductions in neural activity. Negative responses are real signals, not merely artefacts, but they are also produced by vascular "stealing" effects in some circumstances, and their interpretation requires care. What the resolution really is Nominal resolution and effective resolution are different things in fMRI, and confusing them produces overconfident claims. The nominal spatial resolution is the voxel size. Standard whole-brain studies at 3 tesla acquire voxels of two to three millimetres on a side, each containing on the order of a hundred thousand to a million neurons depending on size and region. At 7 tesla, submillimetre voxels are achievable. But the effective resolution is set by the vasculature as much as by the voxel grid. Gradient-echo BOLD is most sensitive to large draining veins, which carry deoxygenated blood away from active tissue and can lie millimetres, sometimes more than a centimetre, downstream from the neurons that caused the response. Robert Turner put the problem memorably in a 2002 paper asking how much cortex a vein can drain. In practice, the peak of a gradient-echo BOLD response is often displaced towards the pial surface of the cortex, where large veins run, rather than located in the grey matter where the neurons are. Several strategies address this. Spin-echo acquisitions are less sensitive to large vessels and weight the signal towards capillaries, at the cost of sensitivity. Techniques that measure cerebral blood volume, such as vascular space occupancy (VASO), developed by Hanzhang Lu and colleagues and adapted for high-field laminar work by Laurentius Huber and others, or that measure blood flow directly with arterial spin labelling, give signals more closely tied to the site of activity. At ultra-high field, these approaches have made it possible to measure responses in different cortical layers, a scale at which input and output layers can in principle be distinguished. For most cognitive studies at conventional resolution, however, the practical lesson is modest: activation maps localise responses to within several millimetres, not to the precise gyral or columnar location that a high-resolution rendering might suggest. Temporal resolution is similarly layered. The TR sets how often the signal is sampled, but the hemodynamic response it samples is itself a smoothed and delayed version of neural events, taking several seconds to rise and many more to return to baseline. Sampling faster does not make the response faster. What faster sampling does provide is better characterisation of the response shape, better separation of physiological noise at cardiac and respiratory frequencies, and more data points per minute, which increases statistical power up to a point. Relative timing differences between conditions within the same region can be measured with a precision of a few hundred milliseconds, because the vascular delay largely cancels, but comparisons of timing between regions are confounded by regional differences in vasculature. A region that "responds first" may simply have faster vessels. The size of the effect It is worth fixing early in the reader's mind how small the BOLD signal is. At 3 tesla, a strong sensory stimulus such as a flickering checkerboard produces signal changes in primary visual cortex on the order of one to a few percent of the baseline intensity. The effects of cognitive manipulations, such as the difference between remembering and forgetting a word or between two types of decision, are typically a fraction of a percent. The noise in a single voxel's time series, arising from thermal noise in the receiver, physiological fluctuations, motion and slow scanner drift, is commonly of comparable or larger magnitude. The signal also scales with field strength. Higher fields increase both the raw signal-to-noise ratio and the size of the BOLD effect, which is why 7 tesla scanners have become important for high-resolution work. But physiological noise also scales with field strength, and above a certain point it, rather than thermal noise, limits the precision of a measurement. At standard resolutions, gains from moving to higher field are therefore smaller than the raw physics would suggest, while problems such as susceptibility dropout and inhomogeneous radiofrequency transmission grow worse. The design implication follows directly. Because any single observation is dominated by noise, fMRI depends on repetition: many trials per condition within a run, several runs per session, and many participants per study. The statistical power of an experiment is determined largely by how many independent observations of the contrast of interest it contains and how cleanly the design separates them from noise and from one another. Every chapter that follows returns, in one form or another, to this point. Acquisition choices are design choices Because the signal is small and the vascular smoothing is fixed, the parameters of the acquisition belong to the experimental design rather than to the radiographer. Four choices recur in almost every study, and each trades one good against another. The first is voxel size. Smaller voxels reduce partial-volume mixing of grey matter with white matter and cerebrospinal fluid and allow finer localisation, but signal-to-noise ratio falls roughly in proportion to voxel volume, so halving the edge length of a voxel costs a factor of eight in raw signal. For a study asking which of two adjacent visual areas responds to a stimulus, the trade is worthwhile. For a study asking whether a large prefrontal network is engaged by a demanding task, it usually is not, because the data will be smoothed and averaged across people in any case. The second is coverage. Whole-brain acquisition is the default for exploratory cognitive studies, but a researcher with a strong hypothesis about a small structure, such as the hippocampus, the amygdala or a brainstem nucleus, may gain far more from a restricted slab acquired at higher resolution or shorter TR. The price is that effects outside the slab cannot be seen, and any claim that a region was "not involved" becomes meaningless for regions not imaged. The third is the repetition time. Shorter TRs provide more samples, better separation of cardiac and respiratory fluctuations, and finer characterisation of response timing. But multiband acceleration introduces its own costs: reduced signal per image because magnetisation has less time to recover, potential leakage of signal between simultaneously excited slices, and larger data files. Several groups have reported that very high acceleration factors can reduce sensitivity in deep structures. Moderate acceleration, giving TRs between roughly 0.7 and 1.5 seconds at around two-millimetre resolution, is a common compromise in current practice, and it is the regime adopted by large projects such as the Human Connectome Project, which used a TR of 0.72 seconds. The fourth is echo time and the number of echoes. A single echo time tuned to average grey-matter T2* is standard. Acquiring several echoes after each excitation, a technique that has become practical with modern acceleration, allows the BOLD component, which scales with echo time, to be separated from other signal changes that do not. Chapter 6 returns to multi-echo acquisition as a tool against motion and physiological noise; here the point is that it must be chosen at the design stage, because no analysis can recover echoes that were never acquired. None of these choices has a universally correct answer. What makes a choice right is its fit to the question. A study built to detect a distributed cognitive effect across a large sample is better served by conventional voxels and robust coverage; a study built to resolve fine spatial structure in a few participants is better served by high field, small voxels and long sessions. The error is to choose parameters by habit or by copying the previous study in the laboratory, and then to ask questions of the data that the parameters could not support. What BOLD cannot tell you on its own Some limits are intrinsic. Because BOLD is a relative signal with no calibrated zero, fMRI cannot measure absolute levels of neural activity. It can only compare one condition with another, or one time with another. The meaning of any activation therefore depends entirely on the comparison that produced it, which is why the choice of control condition, discussed in Chapter 3, is not a technicality but the substance of the inference. Techniques exist to calibrate the BOLD signal against a hypercapnic or hyperoxic challenge and estimate changes in oxygen metabolism, following the approach introduced by Timothy Davis, Kenneth Kwong and colleagues in 1998, but they are demanding and rarely used in cognitive studies. Because the hemodynamic response is shaped by vasculature, comparisons between groups whose vasculature differs are hazardous. Older adults, people with cerebrovascular disease, people taking vasoactive drugs, and children at different developmental stages may all show different BOLD responses to identical neural activity. A finding that older adults "activate" a region less could reflect less neural engagement, altered coupling, or both. Designs that include a control task with a known neural response, against which the experimental response can be compared within each person, offer partial protection. Finally, BOLD reports correlation, not causation. A region whose signal rises during a task is involved in some way with the task, or with some process the task incidentally recruits. Whether the region is necessary for the task is a separate question, answerable only by lesion studies, brain stimulation or pharmacological manipulation. Russell Poldrack's 2006 critique of "reverse inference", reasoning from the activation of a region back to the engagement of a particular mental process, remains essential reading here: most regions respond during many different tasks, so observing activation in a region is weak evidence that any specific process was engaged. None of this diminishes what fMRI can do. It remains the only widely available method that measures activity across the entire human brain, non-invasively, with a spatial precision of a few millimetres. But it does frame the task of the chapters that follow. Every experimental design is a bet about how a slow, indirect, noisy vascular signal will carry information about the mental process under study. The better that bet is informed by the physiology, the better the odds. Chapter 2. The Hemodynamic Response Function If a participant sees a brief flash of light, neurons in primary visual cortex respond within tens of milliseconds and fall silent again within a few hundred. The BOLD signal in the same region does something quite different. For a second or two it barely changes. Then it rises, reaching a peak around five or six seconds after the flash. It falls back over the next several seconds, dips below its starting level, and returns to baseline only after fifteen to twenty seconds or more. This stereotyped time course, the response of the vascular system to a brief burst of neural activity, is the hemodynamic response function, usually shortened to HRF. Nearly every analysis of task fMRI treats it as the lens through which neural events are seen, and nearly every design decision in the following chapters is ultimately a decision about how to work with, or around, its sluggishness. The shape and its parts The canonical HRF has three features that experimenters should be able to sketch from memory. The first is a delay: the response takes one to two seconds to begin rising, reflecting the time needed for the neurovascular signalling cascade to dilate vessels and for fresh blood to arrive. The second is a peak, typically between four and six seconds after a brief event, with a width at half maximum of several seconds. The third is the post-stimulus undershoot, a period during which the signal falls below baseline, often lasting ten seconds or more, before slowly recovering. The undershoot has attracted its own literature. The balloon model, proposed by Richard Buxton, Eric Wong and Lawrence Frank in 1998, explained it mechanically: venous blood volume, having expanded during the response, returns to baseline more slowly than blood flow, so for a time there is an excess of deoxygenated blood in the voxel. Later measurements suggested that a sustained elevation of oxygen metabolism after stimulation also contributes. For analysis the mechanism is secondary; what matters is that the undershoot exists, that it varies in size across regions and people, and that ignoring it produces models that fit the data less well. Some studies, particularly at high field, have reported a small, brief "initial dip" in the signal during the first second or two, attributed to an early rise in oxygen consumption before flow increases. The dip has been more consistently observed in optical imaging of animal cortex than in human fMRI, where it is small, inconsistent and practically irrelevant to standard analyses. Researchers have formalised the canonical shape in several ways. The most widely used, the default in the SPM software developed at the Wellcome Centre for Human Neuroimaging in London, is a "double gamma" function: one gamma probability density describing the positive response, minus a smaller, later gamma describing the undershoot. FSL offers a closely related double-gamma default, and AFNI provides several options, including a single gamma variate after Mark Cohen's 1997 parameterisation and SPM-style double-gamma functions. The SPM defaults are worth knowing precisely because so many published analyses have used them, and they are set out in Table 1. Table 1. Default parameters of the SPM canonical hemodynamic response function. Parameter Default value What it controls Delay of response 6 s Time to peak of the positive gamma Delay of undershoot 16 s Time to peak of the negative gamma Dispersion of response 1 Width of the positive lobe Dispersion of undershoot 1 Width of the undershoot Ratio of response to undershoot 6 Relative amplitude of peak and undershoot Kernel length 32 s Duration over which the function is defined Source: SPM software, function spm_hrf, default parameter vector. These values describe a response peaking roughly five seconds after a brief event, since the peak of the combined function falls slightly earlier than the nominal delay of the positive gamma, and an undershoot whose depth is about one sixth of the peak. They are an average, not a law. Linearity and time-invariance The HRF would be of limited use if it described only the response to a single isolated flash. Its power comes from an assumption: that the BOLD system behaves, approximately, as a linear time-invariant system. Linearity means two things. Scaling the neural input scales the response proportionally, and the response to two events is the sum of the responses to each event on its own. Time-invariance means that an event produces the same response whenever it occurs. If both hold, the predicted BOLD time course for any sequence of events can be computed by convolution: each event is replaced by a copy of the HRF, scaled by the event's strength and placed at its onset, and the copies are summed. A sustained block of stimulation lasting twenty seconds is, on this view, simply twenty seconds' worth of overlapping brief responses, which add up to a response that rises, plateaus and falls. The assumption was tested directly by Geoffrey Boynton, Stephen Engel, Gary Glover and David Heeger in a 1996 study of human primary visual cortex. They varied both the contrast and the duration of visual stimuli and found that responses to longer stimuli could be predicted well from responses to shorter ones, at least for durations of several seconds and above. Anders Dale and Randy Buckner showed in 1997 that responses to visual stimuli presented as little as two seconds apart summed in a roughly linear way, a finding that underpinned the rapid event-related designs discussed in Chapter 4. The linearity is approximate, and it breaks down in predictable ways. When events are very brief or very close together, the response to the second is often smaller than linearity predicts, a refractory effect documented in several studies at intervals under about two seconds, including work by Scott Huettel and Gregory McCarthy in 2000. Karl Friston and colleagues modelled such interactions in 1998 using Volterra series, which add terms capturing how the response to one event depends on recent history. Part of the nonlinearity is neural, as repeated stimuli produce adapted neural responses, and part is vascular, as a vessel that is already dilated cannot dilate proportionally further. The balloon model and its extensions capture some of the vascular component. For most practical purposes the linear model works well enough that it remains the foundation of standard analysis. Two cautions follow. Designs that pack events densely, with intervals well under two seconds, will produce responses that the linear model systematically overpredicts, and the resulting errors can bias comparisons between conditions presented at different rates. And when a study deliberately manipulates repetition, as in adaptation designs that rely on reduced responses to repeated stimuli, the vascular nonlinearity must be distinguished from the neural adaptation the study is trying to measure. Variability across regions, people and populations The canonical HRF is a convenient fiction, and the size of its fictionality has been measured. Geoffrey Aguirre, Eric Zarahn and Mark D'Esposito reported in 1998 that the shape of the response in motor cortex varied considerably across individuals, with peak times differing by several seconds, while remaining relatively consistent across sessions within the same individual. Daniel Handwerker, John Ollinger and D'Esposito extended this in 2004, showing that response shape varied both across participants and across regions within a participant, and that fitting a single canonical shape to all could reduce sensitivity and bias results. Other sources of variability are systematic. Mark D'Esposito, Leon Deouell and Adam Gazzaley reviewed evidence in 2003 that ageing alters neurovascular coupling, with changes in the amplitude, shape and variability of the HRF in older adults. Infants and young children show responses that differ in timing and sometimes in sign from adult responses, as studies of neonatal sensorimotor cortex have documented. Vascular disease, stroke, anaesthesia, and drugs with vascular effects all alter the response. Even within healthy young adults, responses in subcortical structures and in regions near large vessels can have characteristically different timing from those in neocortex. This variability matters in two distinct ways. It reduces power: if the model predicts a peak at five seconds and the true peak in a region is at seven, the model fits the data less well, the estimated amplitude is biased downward, and real effects may go undetected. More seriously, it can create spurious differences. If two groups differ in HRF timing but not in neural activity, and the analysis assumes the same shape for both, the group whose response better matches the canonical shape will appear to activate more strongly. The same logic applies to two conditions whose responses differ in timing, for example because one involves a longer decision process. A difference in amplitude estimates may then reflect a difference in latency rather than in magnitude. Flexible models: derivatives, basis sets and FIR Three broad strategies address response variability, and each trades robustness against flexibility. The first adds derivatives to the canonical shape. Friston and colleagues proposed in 1998 that the canonical HRF be supplemented with its temporal derivative, which captures small shifts in timing, and optionally its dispersion derivative, which captures changes in width. A response that peaks slightly earlier or later than the canonical shape can then be represented as a weighted combination of the canonical function and its temporal derivative. The approach is economical, adding only one or two regressors per condition, and it absorbs variability that would otherwise inflate the residual error. It is less effective for large shifts, beyond about a second, because the approximation on which it rests holds only for small deviations. A subtle point often missed is that the amplitude estimate for the canonical regressor alone no longer fully represents the response when the derivative carries substantial weight; Vince Calhoun and colleagues proposed in 2004 a way to combine the two estimates into a single amplitude measure for group analysis. The second strategy uses a more general basis set, a small family of functions capable of representing a range of shapes. Gamma functions at several delays, cosine sets and the "informed basis set" of canonical plus derivatives are all examples. Such sets fit a wider range of responses, at the price of estimating more parameters and complicating group-level inference, because the question "is there an effect?" becomes a question about several parameters jointly. The third strategy abandons shape assumptions altogether. A finite impulse response (FIR) model represents the response to each condition as a series of independent parameters, one for each time bin following an event: the average signal at zero to two seconds, at two to four seconds, and so on, out to twenty or thirty seconds. The fitted parameters trace out the response shape directly, as estimated from the data. This is deconvolution in its most explicit form, recovering the impulse response from observed signals in which responses to many events overlap. Gary Glover described the approach and its noise properties in 1999. FIR models make no assumptions about shape, but they estimate many parameters, are correspondingly noisier, and can produce implausible shapes when data are limited. They also require that the design support the estimation, a requirement that shapes event-related design, as Chapter 4 explains. Martin Lindquist and colleagues compared these approaches systematically in 2009, examining how well each recovered amplitude, latency and width under various kinds of mismatch. Their central finding was that the canonical model is efficient when its assumptions hold but can be seriously biased when they do not, while flexible models are less biased but more variable. They also proposed a model based on inverse logistic functions that performed well across conditions. No single approach dominates. The right choice depends on whether the goal is detection, where a well-chosen canonical model is usually best, or characterisation of the response, where flexible models are necessary. Duration, reaction time and the event model Before any HRF is convolved with anything, the analyst must decide what the neural "event" is. The choice is less obvious than it appears, and it interacts with the response model in ways that affect conclusions. The simplest convention models each trial as an instantaneous impulse at stimulus onset, sometimes called a "stick" function. This is reasonable when the neural processing evoked by a trial is brief relative to the HRF, as with a flash of light or a single tone. It is less reasonable for trials in which processing lasts for a variable and substantial time, such as a decision that takes anywhere from half a second to three seconds. Under linearity, a longer period of neural activity produces a larger and slightly later BOLD response, even if the intensity of activity per unit time is identical. A region that simply stays engaged for longer on difficult trials will therefore appear, under a stick model, to respond more strongly to difficult trials. Jack Grinband, Tor Wager, Martin Lindquist, Vincent Ferrera and Joy Hirsch examined this problem in 2008, comparing models that treated each trial as an impulse, as a fixed-duration epoch, or as an epoch whose duration equalled the participant's reaction time on that trial. They argued that the variable-duration model was more sensitive to regions whose activity tracked time on task and made clearer what an amplitude difference between conditions meant. The broader issue, that reaction time differences between conditions can produce apparent activation differences throughout the brain, has since been pressed further. Jeanette Mumford, Russell Poldrack and colleagues argued in 2024 that many reported contrasts between conditions of differing difficulty may be substantially driven by time on task rather than by the cognitive process the contrast was meant to isolate, and that the choice of how to model reaction time can change which regions appear to distinguish the conditions. There is no single correct answer, because it depends on the hypothesis. If the question is whether a region's activity per unit time differs between conditions, reaction-time duration modelling, or including reaction time as a separate regressor, is necessary. If the question is whether a region contributes more in total to one condition than another, the difference in duration may be part of the effect of interest. What is not defensible is making the choice by default and interpreting the result as though it answered the other question. The decision belongs in the design, alongside the choice of task, because the task determines how much reaction times will vary and how strongly they will differ between conditions. What deconvolution can and cannot recover It is tempting to think of deconvolution as undoing the hemodynamic blur and recovering the neural events underneath. It does something more limited. When the design includes events at varied intervals and the response is assumed to be the same for every event of a given type, deconvolution estimates the average response to that type of event, separating it from overlapping responses to neighbouring events. It does not recover the neural time course itself. A response shape estimated by an FIR model is still a hemodynamic response, blurred and delayed; it simply has not been forced into a canonical template. Attempts to go further, estimating the underlying neural activity from BOLD, do exist. Blind deconvolution methods, including approaches by Guo-Rong Wu and colleagues developed for resting-state data, and model-based approaches rooted in the balloon model and dynamic causal modelling, attempt to infer latent neural signals. These are ill-posed problems. Many different neural time courses, filtered through a slow HRF and corrupted by noise, produce nearly identical BOLD signals, so the solutions depend heavily on prior assumptions. They are research tools, not routine corrections. The practical lessons for design follow from what deconvolution requires. Estimating the response to a condition requires that the design sample that response at many different delays relative to its onset, which is achieved either by jittering the intervals between events or by desynchronising event onsets from the TR. It also requires that the responses to different conditions not be locked in a fixed sequence, since two conditions that always occur together in the same order cannot be separated no matter what model is used. The following two chapters turn these requirements into concrete design choices. The shape of the HRF, fixed by physiology, determines which choices are available. Chapter 3. Block Designs and the Logic of Comparison The first human BOLD experiments were block designs. Kwong and colleagues in 1992 alternated periods of darkness with periods of flickering light; Bandettini and colleagues alternated rest with finger tapping. The choice was partly inherited from positron emission tomography, where each scan integrated activity over a minute or more and a condition had to be sustained for that long to be measured at all. But it was also a sound choice on its own terms, and it remains so. A block design, in which a participant performs one kind of task continuously for a period of seconds and then switches to another, is the most statistically efficient way to determine whether a region responds differently in two conditions. Understanding why, and understanding what block designs cannot do, is the foundation for everything else in experimental design. Why blocks are powerful Recall from Chapter 2 that the BOLD response to a sustained period of neural activity is the sum of many overlapping hemodynamic responses. As those responses pile up, the signal rises to a plateau and stays there for as long as the activity continues. Alternating two conditions in blocks therefore produces a large, slow oscillation in the BOLD signal of any region that distinguishes them. The difference between the plateau levels is as large as the response can make it. A useful way to see the advantage is in terms of frequency. The HRF acts as a filter. Rapid fluctuations in neural activity, faster than a few seconds, are smoothed out and barely transmitted to the BOLD signal. Very slow fluctuations pass through the filter, but the fMRI signal at very low frequencies is dominated by noise: scanner drift, slow physiological changes and gradual head movement, which together produce noise whose power rises steeply as frequency falls. An experimental manipulation is detected best when it places most of its variance at frequencies the HRF transmits well and where noise is comparatively low. A block design alternating two conditions every fifteen seconds or so puts almost all of its variance at a single frequency, around one cycle every thirty seconds, which lies close to the band the HRF passes most effectively and well above the worst of the drift. This is why the commonly cited guidance is to keep blocks somewhere between about ten and twenty seconds long, giving a full cycle of perhaps twenty to forty seconds. Much shorter blocks begin to lose power because the HRF cannot follow the alternation and the response never reaches its plateau. Much longer blocks shift the signal to lower frequencies, where it is increasingly confounded with drift and where the high-pass filter applied in preprocessing, with a default cut-off period of 128 seconds in SPM, begins to remove it along with the noise. A design with sixty-second blocks, a total cycle of two minutes, has placed its signal almost exactly where the filter will attenuate it. The analyst must then either use a less aggressive filter, admitting more drift, or accept reduced sensitivity. The same logic warns against designs in which a condition occurs only once or twice per run. Its effect cannot be distinguished from a slow trend. Block designs have a second advantage that is easy to overlook: robustness to errors in the HRF model. Because a sustained block produces a plateau whose shape is dominated by the block's duration rather than by the fine details of the response function, misspecifying the HRF's peak time by a second or two has only a modest effect on the estimated difference between conditions. Event-related designs, which depend on the shape of individual responses, are more sensitive to such errors. The logic of subtraction and its limits A block design detects differences. The question is always what the difference means, and the answer depends entirely on what the two conditions were. The dominant logic is cognitive subtraction. Condition A contains process X plus everything else; condition B contains everything else but not X; the difference in activity between them is attributed to X. The approach has a long lineage. The Dutch physiologist Franciscus Donders proposed in 1868 to measure the duration of mental operations by subtracting reaction times across tasks that differed by a single stage, and the influential PET studies of word processing by Steven Petersen, Peter Fox, Michael Posner, Mark Mintun and Marcus Raichle in 1988 built a hierarchy of tasks, from passive viewing of words to reading them aloud to generating associated verbs, each designed to add one processing stage to the last. Subtraction depends on the assumption of pure insertion: that adding process X to a task leaves every other process unchanged. The assumption is frequently false. Adding a semantic judgement to a word-viewing task may change how attentively the word is perceived, how the participant prepares a response, how aroused and engaged the participant is, and how much the participant covertly rehearses. Karl Friston and colleagues set out the problem in a 1996 paper titled "The trouble with cognitive subtraction", arguing that interactions between the added process and the existing ones are the rule rather than the exception. The consequence is that a subtraction can reveal activity that belongs not to the process of interest but to the change it induces in other processes. There is no way to eliminate this problem completely, but there are several ways to constrain it. The first is to match conditions as closely as possible in everything except the manipulation of interest: stimulus properties, response demands, difficulty, and time on task. A contrast between reading words and viewing a fixation cross differs in so many respects that almost any interpretation of the resulting activation is under-determined. A contrast between reading real words and reading orthographically matched pronounceable non-words, with identical response demands, is far more interpretable. The second is the factorial design. Instead of one contrast, the experimenter crosses two or more factors, for example stimulus type (faces versus houses) with task (judge identity versus judge location). Main effects then average over the levels of the other factor, and the interaction, the extent to which the effect of one factor depends on the level of the other, becomes an explicit object of study. Rather than assuming pure insertion, a factorial design tests it: a significant interaction means insertion was not pure. The third is parametric variation. Rather than contrasting presence and absence of a process, the experimenter varies its demand across several levels, such as memory load of one, two, three and four items, or the number of words per minute to be read. A region involved in the process should show activity that changes systematically with the parameter, while processes common to all levels are held roughly constant. Christian Büchel, Richard Wise, Catherine Mummery, Jean-Baptiste Poline and Karl Friston demonstrated early versions of this approach, modelling nonlinear relationships between the rate of word presentation and regional responses. Parametric designs make it harder, though not impossible, for a confound to masquerade as the effect of interest, because the confound would have to vary with the parameter in the same way. The fourth is conjunction. If a region is thought to be engaged by a process X that is present in several different task pairs, then a region that shows the effect in all pairs, each with different incidental confounds, is more credibly associated with X. Cathy Price and Karl Friston proposed conjunction analysis in 1997. Thomas Nichols, Matthew Brett, Jesper Andersson, Tor Wager and Jean-Baptiste Poline showed in 2005 that the most common implementation tested a weaker hypothesis than researchers usually believed, and advocated a test requiring that every contrast be individually significant, the "minimum statistic compared to the conjunction null". The distinction matters: a conjunction test that is valid for the claim "every contrast shows an effect in this region" requires the stricter procedure. Rest is not nothing A special case of subtraction, and a frequent source of confusion, is the comparison with rest. Many block designs alternate a task with periods of fixation or eyes-closed rest, treating rest as a baseline of zero cognitive activity. It is not. Resting participants think. They recall the past, anticipate the future, mind-wander, monitor their bodies and the environment, and sometimes fall asleep. In the late 1990s, Gordon Shulman and colleagues pooled PET studies and found a consistent set of regions, including the medial prefrontal cortex, posterior cingulate and precuneus, and the lateral parietal cortex, that were more active during passive baseline conditions than during a wide range of goal-directed tasks. Marcus Raichle and colleagues argued in 2001 that these regions formed a "default mode" of brain function, active by default and suspended during externally focused tasks. The medial temporal lobe presents a related problem. Craig Stark and Larry Squire showed in 2001 that activity in the hippocampal region during rest could be higher than during some tasks, so that memory-related activation measured against rest could be obscured or even reversed. Using an active baseline, such as a simple arrow-direction judgement, reduced this problem. The practical point is that "task versus rest" contrasts are dominated by the difference between an externally focused state and an unconstrained one. They are excellent for localising regions that respond to a class of stimulus or task, and they provide a useful manipulation check. They are poor tests of any specific cognitive hypothesis. When rest periods are included, it is often better to treat them as a measure of the signal's floor for modelling purposes, while making the theoretically interesting comparison between two active conditions. The costs of blocking Block designs buy power with a set of specific limitations, and a design that ignores them may produce a clean activation map with an ambiguous meaning. The first limitation is predictability. Within a block, participants know what kind of trial is coming next. This changes their state: they can prepare, adopt strategies, allocate attention, and settle into a rhythm. Some processes of interest, such as responses to surprise, novelty or conflict, are altered or abolished when trials are predictable. A block of "oddball" stimuli, in which every stimulus is the oddball, is no longer an oddball paradigm. The second is that the block average conflates transient and sustained activity. A region that responds briefly to each trial and a region that maintains a sustained state across the block can produce similar block-level effects. Some questions require separating them, which is the rationale for the mixed designs described in Chapter 4. The third is that trials cannot be sorted after the fact. If the question concerns trials that were later remembered versus those later forgotten, or correct versus error trials, or trials on which the participant reported seeing a threshold stimulus versus those on which they did not, the relevant categories are only known after the response, and they are interleaved unpredictably within any block. Block designs cannot separate them without mixing responses across categories. The fourth is habituation and fatigue. Sustained stimulation of one kind can produce neural adaptation within a block, and long runs of repetitive blocks can produce declining arousal and attention over the session. Both reduce effects and can differ across conditions. Counterbalancing condition order across runs and participants, and keeping runs to a moderate length of perhaps five to ten minutes, mitigates both. The fifth is that blocks invite state-based confounds. Because conditions are separated in time and sustained, anything that differs systematically between blocks, including breathing patterns, eye movements, arousal, or head movement, is correlated with the task at the same slow frequency as the effect of interest. A task that makes participants hold their breath, speak aloud or shift their gaze will produce BOLD fluctuations at the block frequency that have nothing to do with neural computation in the regions of interest. Chapter 6 treats these confounds in detail. Here the design lesson is that such confounds should be anticipated, measured and, where possible, balanced across conditions. Localisers and regions of interest Block designs have found a lasting role in functional localisers: short, efficient scans designed to identify a functionally defined region in each participant, which is then interrogated in a separate main experiment. The classic example is the fusiform face area, characterised by Nancy Kanwisher, Josh McDermott and Marvin Chun in 1997 using blocks of faces alternated with blocks of objects. A localiser of this kind, run in a few minutes, identifies a region in each individual's brain with far greater precision than any group-average coordinate can provide, because the exact location of face-selective cortex varies across people by a centimetre or more. The approach has been debated. Karl Friston, Pia Rotshtein, Joy Geng, Philipp Sterzer and Richard Henson published "A critique of functional localisers" in 2006, arguing that separate localiser scans are often less efficient than a factorial design that includes the localising contrast as one of its factors, and that restricting analysis to a localised region risks missing effects elsewhere. Rebecca Saxe, Matthew Brett and Nancy Kanwisher replied the same year in "Divide and conquer", defending localisers on the grounds that they allow hypotheses to be tested in functionally equivalent regions across individuals, that they reduce the multiple-comparisons burden, and that they permit results to be compared across studies. Both sides agree on one principle that is essential to any region-of-interest analysis: the data used to define a region must be independent of the data used to test effects within it. Selecting voxels because they show a strong difference between conditions A and B, and then reporting that those voxels show a strong difference between A and B, is circular. Nikolaus Kriegeskorte, W. Kyle Simmons, Patrick Bellgowan and Chris Baker documented in 2009 how such "double dipping" inflates effect sizes, and Edward Vul, Christine Harris, Piotr Winkielman and Harold Pashler reported the same year that implausibly high correlations between brain activity and personality measures in social neuroscience often arose from this kind of non-independent selection. A separate localiser run, or a split of the main data into independent halves, avoids the problem. A worked planning example Consider how these principles combine in a concrete, if hypothetical, case. A researcher wants to know which regions show greater activity when participants hold a sequence of letters in mind and compare each new letter with the one presented two items earlier, the familiar two-back task, than when they simply watch for a single target letter, the zero-back task. The two conditions share visual input, button presses, and the general demands of attending to a stream of letters. They differ in working-memory maintenance and updating, which is the process of interest, but also in difficulty, error rate, and probably arousal. A first decision concerns block length. With letters presented every two seconds, a block of ten letters lasts twenty seconds, long enough for the BOLD response to approach its plateau and short enough that the alternation between conditions sits comfortably above the high-pass filter. Adding a brief instruction screen at the start of each block, and modelling it as a separate regressor, prevents instruction reading from being absorbed into the task effect. A second decision concerns the baseline. Including short fixation periods between some blocks gives the model an estimate of the signal floor and allows each task to be compared with rest, which is useful as a manipulation check. But the inferential contrast is two-back versus zero-back, not two-back versus rest, because the latter would be dominated by the difference between any task and none. A third decision concerns confounds. Because two-back is harder, participants may make more errors, move more, and breathe differently. Recording responses allows error trials to be examined; recording respiration allows its relation to the block structure to be checked. Including a parametric element, for example a one-back condition as an intermediate level, would let the analysis ask whether activity scales with load across three levels, which is harder for a generic difficulty confound to mimic than a single two-level contrast. A fourth decision concerns order and length. Running the task in two runs of about five minutes each, with condition orders counterbalanced across runs and participants, limits fatigue and prevents any condition from being systematically confounded with time in the scanner. The design that results is unremarkable to look at, and that is the point: every one of its features was chosen to protect a specific inference. When to choose a block design The decision can be put simply. If the primary question is whether, and where, activity differs between two or a few well-matched conditions whose trials can be grouped in advance, and if predictability does not destroy the process of interest, a block design will detect the difference with fewer participants and less scanning time than any alternative. It is especially suited to localisers, to clinical applications such as presurgical mapping of language and motor areas, and to studies in populations who cannot tolerate long sessions. If the question involves the timing or shape of responses, the separation of trial types that are determined only by behaviour, processes that require unpredictability, or the dissection of transient from sustained activity, block designs cannot answer it well, no matter how the data are later analysed. Those questions require event-related or mixed designs, the subject of the next chapter. Hashtags: #FunctionalMagneticResonanceImaging #FMRI #BOLDImaging #BOLDSignal #ExperimentalParadigmDesign #TaskFMRI #BlockDesign #EventRelatedDesign #MixedDesign #HemodynamicResponseFunction #HRF #GeneralLinearModel #DesignMatrix #ContrastAnalysis #NeurovascularCoupling #HeadMotionCorrection #PhysiologicalNoise #SpatialNormalization #StatisticalParametricMapping #MultipleComparisons #ClusterInference #PermutationTesting #FunctionalLocalizers #NeuroimagingMethods #FutureOfFMRI
- Community-Based Participatory Research (Collaborative Methodologies for Social Impact)
Download the Book (PDF): Introduction In the late 1980s residents of West Harlem began organizing against the environmental burdens their neighborhood had been asked to carry: a sewage treatment plant on the Hudson waterfront, and a disproportionate share of Manhattan's diesel bus depots. Northern Manhattan's children were being hospitalized for asthma at rates far above the city average. Residents did not need a study to tell them the air was bad; they could smell it and they could see their children's inhalers. What they needed was evidence that would count in rooms where they had no seat. The organization that grew out of that fight, West Harlem Environmental Action, founded in 1988 and now called WE ACT for Environmental Justice, entered into a long partnership with researchers at Columbia University's school of public health. In one early pilot study, published by Patrick Kinney, Peggy Shepard and colleagues in 2000, community members and scientists together measured fine particles and diesel exhaust on Harlem sidewalks and found that concentrations tracked the traffic of trucks and buses street by street. Measurements of that kind became part of the case that pushed New York's transit authority toward cleaner buses and depot reforms. The story is often told as a success of community-based participatory research, and it was. But the most instructive thing about it is not the air monitors. It is the order of events. The community named the problem before any researcher arrived. The community decided that measurement was the tool it needed. The researchers were invited into a fight that was already under way, and the value of their contribution depended on whether they could bring the technical capacity of a university without taking the fight away from the people who had started it. That order of events is the subject of this book. Community-based participatory research, usually shortened to CBPR, is a family of approaches in which people affected by a problem take part as partners, not subjects, in studying it. In the definition that has shaped the field in public health, it is a collaborative approach that equitably involves community members, organizational representatives and researchers in all aspects of the research process, with each partner contributing unique strengths and shared responsibilities. The key words in that sentence are "equitably" and "all aspects." Many projects involve communities. Far fewer involve them equitably, and fewer still involve them in all aspects, from the choice of question to the ownership of what was learned. The argument The argument of this book is that participation is real only to the extent that it redistributes control over decisions that carry consequences. Research is a sequence of such decisions. Someone decides what question is worth asking and what counts as an answer. Someone decides who collects the data and on what terms. Someone decides what the data mean, and whose interpretation prevails when two readings conflict. Someone's name goes on the paper. Someone owns the samples, the data files and any invention that follows. Someone decides what happens to all of it when the grant ends. At each of these points, control can move toward the community or quietly return to the institution. A project is participatory at the points where it moved, and conventional at the points where it did not, whatever its funding application says. This way of seeing the work has a practical advantage. Much of the literature on participatory research describes principles: trust, respect, mutual benefit, colearning. The principles are sound, but principles are cheap to endorse and hard to audit. Decisions, by contrast, leave traces. One can ask who was in the room when the research question was fixed, who signed the data use agreement, who is listed as first author, and who holds the keys to the database. A partnership that cannot answer those questions clearly has not yet decided how power will be shared, and in the absence of a decision the default is the institution, because the institution holds the money, the ethics approval, the publication channels and the legal personality that contracts require. The argument also explains why CBPR is harder than its admirers sometimes admit. Redistributing control has costs. It takes time that grant cycles do not budget for. It gives community partners the ability to say no, including to questions a researcher cares about and to publications a researcher needs for promotion. It exposes academic credit systems, built around individual authorship and priority, to claims they were not designed to handle. And it forces uncomfortable conversations about ownership in a legal environment that generally assumes the institution owns what its employees produce. A partnership that has never felt any of these costs has probably not shared much control. What the book covers and what it leaves out The chapters follow the life of a research project, because the decisions that matter arrive in roughly that order. The first chapter traces where participatory research came from, since its competing traditions still shape what practitioners mean by it. The second examines power directly: the structures, agreements and habits through which partnerships allocate control, and the ways control leaks back to the institution. The next three chapters take the core research tasks in turn: framing the question, collecting data, and interpreting findings. The sixth and seventh chapters take up the two issues where participatory ideals most often collide with institutional rules: authorship and credit, and ownership of data and intellectual property. The eighth asks what a partnership leaves behind when the funding stops, which is where many communities judge whether the whole exercise was worth their time. The book draws its examples mainly from health and environmental research, because that is where CBPR has been most fully developed, evaluated and argued over, and from research with Indigenous peoples, because Indigenous nations have done more than anyone to turn the ethics of research into enforceable rules. The principles transfer to education, urban planning, criminal justice, disability research and development studies, and examples from those fields appear where they sharpen a point. The book does not attempt a manual of specific data collection techniques, and it does not survey the large literature on citizen science in ecology and astronomy, where volunteers typically contribute observations to projects designed by others. That arrangement has great value, but it answers a different question from the one pursued here. A note on words "Community" is an unstable word, and the instability matters. A community can be a neighborhood, a tribal nation with its own government, a group of people sharing a diagnosis, a workforce, or a population defined by a researcher's sampling frame. These are very different partners. A sovereign nation can pass laws governing research on its lands; a neighborhood association cannot. A patient advocacy group may have paid staff and a policy agenda; a loosely connected group of residents may have neither. When this book speaks of the community partner, it means whichever people and organizations have a legitimate claim to represent those most affected by the research, and it treats the question of who holds that claim as part of the work, not a preliminary to it. Similarly, "researcher" here usually means an academic or institutionally employed researcher, and "institution" means the university, hospital, agency or firm that employs them, holds the grant and signs the contracts. The division is a convenience. Many community members are researchers by training, and many academic researchers come from the communities they study. But the division tracks something real: the unequal distribution of the resources that research requires, and the unequal exposure to the harms it can cause. The book is written for researchers who want to do this work well, for community organizations deciding whether and how to partner with a university, for funders and ethics committees trying to judge whether a proposal's claims to participation are genuine, and for students entering fields where community engagement is now expected. Each will find that the questions are the same from different sides of the table. The answers, when they are good, are agreed between those sides and written down. Chapter 1. Two Lineages of Participation Community-based participatory research did not begin as a method. It began as a set of arguments about who is entitled to produce knowledge, and those arguments came from two directions that have never been fully reconciled. One tradition treats participation as a way to make research work better: more relevant questions, better recruitment, more valid measures, findings that are actually used. The other treats participation as a way to shift power: research as a tool with which oppressed people analyze and change their own situation. The first tradition asks how communities can help research. The second asks how research can help communities, and whether research controlled by outsiders can ever do so. Most contemporary CBPR sits somewhere between them, and many of the disputes inside partnerships are really disputes about which lineage the project belongs to. The northern tradition: action research and the usefulness of involvement The first lineage is usually traced to Kurt Lewin, the German-American social psychologist who coined the phrase "action research" in the 1940s. In a 1946 paper in the Journal of Social Issues, "Action Research and Minority Problems," Lewin argued that research which produced nothing but books would not suffice, and proposed a cycle of planning, action and fact-finding about the results of action. His concern was practical. He worked with community relations bodies trying to reduce intergroup prejudice, and he observed that the people charged with acting on a problem learned more, and acted more effectively, when they were involved in studying it. Lewin's framing proved durable because it made participation an instrument of effectiveness. In organizational development, education and later in health services research, action research came to mean practitioners studying their own practice in iterative cycles, often with an outside facilitator. Teachers investigated their own classrooms; nurses studied their own wards. The approach was participatory in that the people closest to the work did the inquiry, but it did not necessarily challenge who held power in the organization. A hospital could run an action research project that improved discharge planning without asking whether patients should have a say in what counted as a good discharge. This lineage has much to recommend it. It produced a large body of practical knowledge about how to run iterative inquiry, how to combine reflection with change, and how to make findings usable. It also produced one of the central claims of modern CBPR: that involving the people affected by a problem can improve the science itself, by surfacing variables outsiders miss, by designing instruments that people actually understand, and by recruiting participants who would not answer a stranger's knock. When funders today justify community engagement, they usually do so in these terms. The southern tradition: knowledge as power The second lineage emerged in Latin America, Africa and South Asia in the 1960s and 1970s, and its founding texts read very differently. Paulo Freire's Pedagogy of the Oppressed, first published in Portuguese in 1968 and in English in 1970, argued that education could either domesticate people or free them, and that liberating education began with learners naming their own world. Freire's literacy circles in northeastern Brazil started from "generative themes" drawn from peasants' own lives, not from textbooks written elsewhere. The method assumed that poor and illiterate people already possessed knowledge about their circumstances, and that the educator's job was to help them analyze it critically, not to deposit correct information into them. The Colombian sociologist Orlando Fals-Borda carried a similar conviction into research. Working with peasant movements on Colombia's Atlantic coast in the 1970s, he developed what came to be called participatory action research, in which the separation between researcher and researched was deliberately broken down and the results were returned to the communities in forms they could use, including popular publications and illustrated histories. Fals-Borda and Muhammad Anisur Rahman later collected accounts of this work from several continents in Action and Knowledge: Breaking the Monopoly with Participatory Action-Research (1991). The subtitle states the thesis. Academic research held a monopoly on legitimate knowledge, and that monopoly was one of the ways in which the powerful stayed powerful. In Tanzania, Budd Hall and colleagues working in adult education in the early 1970s used the term "participatory research" for similar work, and Hall went on to help build international networks around it. In India, Rajesh Tandon founded the Society for Participatory Research in Asia (PRIA) in 1982. Robert Chambers, at the Institute of Development Studies in Sussex, popularized participatory rural appraisal in the 1980s and 1990s, a family of techniques such as community mapping, seasonal calendars and wealth ranking through which villagers could analyze their own conditions. Chambers's book Whose Reality Counts? Putting the First Last (1997) asked the question in its title of development professionals who had been designing programs from capital cities. This tradition does not regard participation primarily as a route to better data. It regards the conventional research relationship, in which an outsider extracts information and takes it away to be analyzed and published elsewhere, as itself a form of domination. The remedy is not to involve communities more skillfully in the outsider's project but to make the project theirs. Other currents Two further currents fed into contemporary CBPR and gave it some of its sharpest tools. Feminist researchers from the 1970s onward challenged the claim that good research required detachment, argued that the researcher's position shaped what she could see, and developed methods such as collaborative interviewing that treated participants as knowers. Their emphasis on reflexivity, the discipline of examining how one's own identity and interests shape the research, is now standard in participatory work. Indigenous scholars made the most direct challenge. Linda Tuhiwai Smith's Decolonizing Methodologies: Research and Indigenous Peoples, first published in 1999, opened with the observation that "research" is probably one of the dirtiest words in the Indigenous world's vocabulary. Her book documented how research had served colonial projects: classifying, measuring and displaying Indigenous peoples, removing their remains and artifacts, and producing knowledge that justified dispossession. Smith, a Māori scholar, set out an agenda in which Indigenous communities would define research priorities themselves, grounded in their own values and protocols. In Aotearoa New Zealand this became Kaupapa Māori research; in Canada and the United States, tribal nations and First Nations organizations developed their own research codes and review boards. Indigenous research ethics now supplies much of the most concrete guidance available on the questions this book treats as central: ownership, control and benefit. Disability rights activism contributed the slogan that many participatory researchers now use, "Nothing about us without us," which James Charlton took as the title of his 1998 book on disability oppression. The slogan compresses the core claim of the power-sharing tradition into five words. The formation of CBPR in public health The term "community-based participatory research" took hold in North American public health in the 1990s. A key text was a 1998 review in the Annual Review of Public Health by Barbara Israel, Amy Schulz, Edith Parker and Adam Becker, "Review of Community-Based Research: Assessing Partnership Approaches to Improve Public Health." It set out a list of principles that has been quoted ever since: CBPR recognizes community as a unit of identity; builds on strengths and resources within the community; facilitates collaborative, equitable partnership in all phases of research; promotes colearning and capacity building among all partners; integrates knowledge and action for mutual benefit; addresses health from positive and ecological perspectives; disseminates findings and knowledge gained to all partners; and involves a long-term commitment by all partners. Later versions added attention to cyclical and iterative processes and to the social inequalities that shape health. Israel and colleagues were writing from experience. The Detroit Community-Academic Urban Research Center, established in 1995 with funding from the US Centers for Disease Control and Prevention, brought together the University of Michigan schools of public health, community organizations on Detroit's east side and the city's health department. Its board, which brought representatives of community-based organizations together with the university, the health department and a health system, reviewed research proposals, and the partners adopted written CBPR principles to govern their joint work. The center became one of the most studied examples of long-term CBPR infrastructure, and the principles it developed influenced many later partnerships. The field grew quickly. Meredith Minkler and Nina Wallerstein edited Community-Based Participatory Research for Health, first published in 2003 and now in its third edition (2018, with Bonnie Duran and John Oetzel as co-editors), which became the standard reference. Community-Campus Partnerships for Health, a network founded in 1996, promoted partnership principles across universities and community organizations. In 2004 the US Agency for Healthcare Research and Quality commissioned a systematic review, led by Meera Viswanathan, of the evidence on CBPR, which concluded that the approach was promising but that the evidence base was thin and inconsistently reported. The journal Progress in Community Health Partnerships was launched in 2007 to give partnerships a venue for publishing their work. By the 2010s US federal funders, including the National Institutes of Health through its Clinical and Translational Science Awards and the Patient-Centered Outcomes Research Institute, created by the 2010 Affordable Care Act, were requiring or rewarding community and patient engagement. Parallel streams: patient involvement and participatory design The same questions surfaced in fields that rarely cited Freire or Fals-Borda. In the United Kingdom, patient and public involvement in health research became an organized expectation from the mid-1990s. The Department of Health established a body in 1996 that later became INVOLVE, and the National Institute for Health Research, created in 2006, built involvement into its funding requirements, so that grant applications must explain how patients and the public shaped the proposal and will shape the study. The language is different from CBPR; the unit is usually a patient group or a panel of lay advisers rather than a geographically bounded community, and involvement often stops short of shared authority. But the underlying logic, that people with lived experience of a condition know things about it that researchers do not, is the same. In Scandinavia, participatory design emerged in the 1970s from collaborations between computer scientists and trade unions. Projects such as UTOPIA, a Swedish and Danish collaboration with graphic workers in the early 1980s, set out to design workplace technologies with the workers who would use them, on the explicit premise that technology choices were also choices about power in the workplace. The tradition survives in human-computer interaction and in the design of public services, and it has contributed practical techniques for involving non-specialists in technical decisions: prototypes people can handle, scenarios they can argue about, and workshops in which the specialist's role is to make options visible rather than to choose among them. These streams matter here for two reasons. First, they show that the questions CBPR raises are not peculiar to public health or to marginalized neighborhoods; they arise whenever the people who produce knowledge and the people who live with its consequences are different people. Second, they illustrate how the same vocabulary can cover very different allocations of power. A patient panel that comments on a lay summary and a patient panel that holds a veto over the study design are both described as "involvement." Only the second has shifted control. What the evidence shows Because participatory research is often defended by its effects, it is fair to ask what those effects are. The honest answer is that the evidence is substantial but uneven, and that it supports some claims more strongly than others. The 2004 review for the Agency for Healthcare Research and Quality found that community involvement was associated with better recruitment and retention and with interventions better matched to local conditions, but that few studies were designed to test whether participation itself caused these results. Margaret Cargo and Shawna Mercer, reviewing the field in the Annual Review of Public Health in 2008, argued that the value of participatory research lay along three dimensions: translating research into practice, advancing self-determination, and pursuing justice. They noted that the second and third were rarely measured at all. A more useful approach to the evidence came from a realist review led by Justin Jagosh and published in the Milbank Quarterly in 2012. Instead of asking whether participatory research "works," the team asked what mechanisms produce its benefits and under what conditions. Examining a set of long-running partnerships, they identified partnership synergy, the combination of perspectives, skills and resources that no partner could achieve alone, as the central mechanism, and they found that its effects were cumulative: early trust made later co-governance possible, which produced culturally appropriate interventions, which in turn deepened trust. They described "ripple effects" that extended beyond the original project, including new partnerships, community capacity and systemic changes. A follow-up realist evaluation published in BMC Public Health in 2015 elaborated how conflict, when handled well, could strengthen rather than break these partnerships. The Engage for Equity study, led by Nina Wallerstein and John Oetzel with colleagues at the University of New Mexico and elsewhere, took the quantitative route. Surveying academic and community partners in roughly two hundred federally funded community-engaged projects in the United States, the team examined associations between partnership practices, such as shared decision-making and community influence over resources, and outcomes ranging from the quality of the research to changes in policy and community capacity. Their findings, reported in a series of papers including a 2020 special section of Health Education & Behavior, supported the proposition that the practices through which power is shared are associated with better outcomes, not just with partner satisfaction. None of this proves that every participatory project produces better science. Partnerships can fail, and a study can be highly participatory and badly designed. But the weight of the evidence runs in a consistent direction: the benefits of participation come less from the presence of community members than from what they are able to decide. That finding is the empirical counterpart of the argument this book makes on ethical grounds. Why the lineages still matter Institutional adoption brought CBPR money, legitimacy and training programs. It also brought a drift toward the northern lineage. When participation is justified by what it does for research quality and recruitment, it becomes easy to measure success by research outcomes and to treat the community as a stakeholder to be consulted rather than a partner with authority. Nina Wallerstein and Bonnie Duran, writing in the American Journal of Public Health in 2010, described CBPR as sitting on a continuum from a utilitarian pole, concerned with making research more effective, to an emancipatory pole, concerned with social transformation, and argued that the field's contribution to health equity depended on not losing the second pole. The two lineages give different answers to almost every concrete question a partnership faces. Who should decide the research question? The utilitarian answer is that researchers should, with community input to make it relevant; the emancipatory answer is that the community should, with researchers helping to make it answerable. Who owns the data? The utilitarian answer defers to institutional policy; the emancipatory answer treats the community's claim as primary. What counts as success? A publication and a funded follow-up grant, or a change in the conditions people live under? Andrea Cornwall and Rachel Jewkes, in a much-cited 1995 paper in Social Science & Medicine, "What Is Participatory Research?", made the point that the key difference between participatory and conventional research lies not in methods but in the location of power in the research process. A focus group is not participatory because it is a focus group; a survey is not extractive because it is a survey. What matters is who decided to use it, who designed it, who holds the results and who decides what they mean. This is the thread the rest of this book follows. A warning came from the same tradition. In 2001 Bill Cooke and Uma Kothari edited a collection provocatively titled Participation: The New Tyranny?, arguing that participatory methods in development had become a ritual that legitimized decisions already made, gave outsiders access to local knowledge without transferring control, and could override local power dynamics in the name of "the community." The critique was not that participation was bad but that the word had become detached from any redistribution of power. It is still the best test to apply to any project that calls itself participatory: what, specifically, can the community decide that it could not decide before? Most partnerships will never be purely emancipatory. They operate inside universities and funding systems that impose deadlines, require principal investigators, demand ethics approval and reward publication. The realistic goal is not purity but honesty: knowing which decisions are genuinely shared, which are delegated to the community, and which remain with the institution, and saying so openly. The next chapter examines how partnerships make those allocations, and how they go wrong. Chapter 2. Power Is the Method Every research partnership has a governance structure, whether or not anyone designed it. If no one decides how decisions will be made, they will be made by whoever controls the budget, signs the ethics application and answers to the funder. In most projects that is the academic principal investigator. This is not usually a matter of bad faith. It is a matter of defaults: universities are built to route authority through principal investigators, and the path of least resistance runs straight to them. Sharing power therefore requires deliberate design, and the design has to be specific enough to survive the moment when partners disagree. The ladder and its limits The most widely cited account of participation remains Sherry Arnstein's "A Ladder of Citizen Participation," published in the Journal of the American Institute of Planners in 1969. Arnstein, who had worked on federal urban renewal and anti-poverty programs, described eight rungs grouped into three bands. At the bottom were manipulation and therapy, which she called nonparticipation: programs designed to educate or "cure" participants rather than let them influence anything. In the middle were informing, consultation and placation, which she called degrees of tokenism: citizens could hear and be heard but had no assurance their views would be acted on. At the top were partnership, delegated power and citizen control, the degrees of citizen power. Her central point was blunt. Participation without redistribution of power is an empty and frustrating process for the powerless, because it allows those in power to claim that all sides were considered while ensuring that only some benefit. Arnstein's ladder has been criticized as too linear. Real partnerships do not sit on a single rung; a community may control recruitment while the university controls analysis, or hold a veto over publication while having little say in the budget. The International Association for Public Participation's spectrum, which runs from inform through consult, involve and collaborate to empower, is a less judgmental version often used in public agencies. Other writers have argued that the top of the ladder is not always the right goal: some communities do not want to run a research project, only to ensure it is done on their terms. The value of the ladder is not that it tells a partnership where to stand, but that it forces the question of where it actually stands on each decision, as opposed to where it says it stands. A practical way to use it is to break a project into its main decisions and ask of each who proposed, who could object, and who had the final word. The pattern that emerges is usually more informative than any overall characterization. Table 1 sets out how several common forms of community involvement tend to allocate a few of the decisions that matter most. Table 1. How common forms of community involvement allocate key decisions. Form of involvement Research question Budget Data custody Publication Community advisory board Researcher sets; board comments Researcher controls Institution Researcher decides Community-engaged study Researcher sets with input Small subcontract to partner Institution Partner reviews drafts Equitable CBPR partnership Jointly negotiated Shared, written allocation Joint agreement Joint approval and authorship Community-led research Community sets Community holds grant Community Community decides The table simplifies, but it makes one pattern plain. The forms differ less in how often community members attend meetings than in who holds money, data and the right to decide what gets published. A community advisory board that meets monthly and is warmly thanked in the acknowledgments may exercise less power than a partner organization that meets rarely but holds a third of the budget and a veto over data release. Structures that share control Partnerships have developed a set of tools for making power-sharing concrete. None is sufficient alone, and each can be hollowed out, but together they turn principles into obligations. The first is a written partnership agreement, sometimes called a memorandum of understanding or a research agreement. At its best it states the project's goals, the roles and responsibilities of each partner, how decisions will be made and disputes resolved, how money will be allocated, who will own and have access to data, how findings will be reviewed before publication, how authorship will be decided, and what happens if a partner withdraws. Many partnerships derive theirs from published models; the Detroit Community-Academic Urban Research Center's principles and the research agreement templates developed by Indigenous health organizations are frequently adapted. The value of such an agreement lies less in its legal force, which is often limited, than in the conversations required to write it. A partnership that has argued about data ownership before any data exist has a much better chance of resolving the argument when the data arrive. The second is a governance body with real authority. The difference between an advisory board and a steering committee is whether its decisions bind. Some partnerships give a community-majority steering committee formal approval over research questions, instruments and publications; others require consensus of all partner organizations on major decisions. Voting rules matter. A committee where the university holds as many seats as all community organizations combined, and the principal investigator chairs, will reproduce institutional control even if every vote is unanimous, because the agenda and the information flow run through the chair. The third is money. Budgets are statements of priorities, and in CBPR they are also statements of power. The principle is simple: community partners who do research work should be paid for it, and community organizations should receive a share of the grant sufficient to cover their real costs, including staff time, overhead and the unglamorous work of convening residents. The practice is harder. Universities often apply their full indirect cost rate to the entire grant while passing through only direct costs to subcontracted community partners; payment systems can take months to process invoices, which a small nonprofit cannot absorb; and some funders restrict what can be paid to people without formal credentials. Partnerships that take equity seriously negotiate these details before submitting the grant, and some have persuaded funders to name a community organization as co-principal investigator or as the lead applicant. The fourth is time and process: meeting at times and places community partners can attend, providing food and childcare, translating materials, and allowing enough time for community partners to consult the people they represent before a decision is made. These are sometimes treated as courtesies. They are in fact conditions for shared decision-making. A decision taken at a weekday afternoon meeting on campus, on the basis of a document circulated the night before in academic English, has not been shared regardless of who was invited. Positionality and the invisible allocation of authority Formal structures do not capture everything. Power also moves through expertise, language and social position. When a researcher explains why a proposed survey question would compromise validity, community partners may defer even when their objection was sound, because the vocabulary of methodology carries authority. When a community partner describes what residents will and will not tolerate, researchers may treat it as anecdote rather than evidence. The literature calls attention to this through the concept of positionality: the recognition that each partner's race, class, gender, education, institutional role and relationship to the community shape what they see and how they are heard. Michael Muhammad and colleagues, in a 2015 article in Critical Sociology, drew on interviews with academic and community partners to describe how researchers' identities, including their institutional privilege and, for researchers of color, their complicated position as both insiders and outsiders, affected trust, decision-making and outcomes. Their recommendation was not that researchers should confess their privilege and move on, but that partnerships should build regular, structured reflection into their work so that these dynamics could be named and addressed. A practical form of this is the power analysis, conducted periodically, in which partners ask of recent decisions who raised the issue, whose information was decisive, whose preferences prevailed, and whether anyone felt unable to object. Such an exercise is uncomfortable, and it works best when a trusted facilitator not employed by the university runs it. It is also one of the few ways to detect the gradual drift of authority back toward the institution, which rarely happens through any single decision. Who speaks for the community? Power-sharing presupposes a partner with whom to share it, and identifying that partner is itself an exercise of power. Researchers often begin with the organizations they already know, or those with the capacity to manage a subcontract: established nonprofits, health clinics, churches with paid staff. These organizations may be deeply rooted, or they may represent the more organized and resourced part of a community while speaking for the rest. A partnership built entirely with service providers may miss the people those providers serve; a partnership with elected tribal leadership may not capture the views of urban tribal members or of dissenting factions. There is no formula for resolving this, but there are better and worse practices. Better partnerships involve several organizations with different constituencies, create channels for residents who are not affiliated with any organization, and revisit the question as the work evolves. They recognize that communities contain their own inequalities, of gender, age, class, caste, immigration status and disability, and that the most affected people may be those with least voice even within their own community. Cooke and Kothari's warning, that participation can reinforce local power structures by treating the loudest voices as the community's voice, applies with full force here. For sovereign Indigenous nations the question has a clearer answer: the nation's government, through whatever processes it has established, has authority to approve or refuse research involving its citizens and lands. Many tribal nations in the United States have established research review boards or research codes; the Navajo Nation's Human Research Review Board, created in the 1990s, is among the best known. In Canada, the Tri-Council Policy Statement on ethical conduct for research, the joint policy of the federal research agencies, devotes a chapter to research involving First Nations, Inuit and Métis peoples, requiring engagement with the relevant community where research is likely to affect its welfare. These arrangements do not eliminate internal debate, but they locate authority where a partnership can find it. A worked example: Kahnawake The Kahnawake Schools Diabetes Prevention Project, begun in 1994 in the Mohawk community of Kahnawake near Montreal, is one of the most thoroughly documented examples of how governance can be designed to keep authority with a community over decades. The project was initiated in response to high rates of type 2 diabetes in the community and was designed as a school- and community-based prevention program with an evaluation component led jointly by community and academic researchers. Several features of its design repay attention. First, the project adopted a written code of research ethics, first in the mid-1990s and revised later, which set out the obligations of community and academic researchers, the principles of the partnership and the procedures for approving publications and handling data. Second, it established a community advisory board, made up of community members, which was not a sounding board for researchers' plans but a body with authority over the project's direction, including review of research proposals and of manuscripts. Third, the research team included community researchers employed by the project alongside academic investigators, so that expertise in both research methods and community life was held inside the team rather than traded across its boundary. Fourth, the project was structured from the outset as long-term: the intervention and its evaluation continued for years, and the partnership outlived several funding cycles. Researchers who worked with the project, including Ann Macaulay and Margaret Cargo, have written about both its strengths and its tensions: the time required for community review, the occasional disagreements between the advisory board and academic researchers over what to measure and publish, and the effort required to keep a partnership vital as individuals moved on. What made the arrangement durable was not the absence of tension but the fact that the procedures for resolving it had been agreed in advance and were understood to bind everyone, including the academics whose careers depended on publication. The lesson generalizes. When the community's authority is written into the project's founding documents and exercised routinely, it becomes part of how the work is done rather than an exceptional intervention. When it exists only as goodwill, it tends to be exercised only when nothing important is at stake. The hidden costs of sharing It is worth being candid about what power-sharing costs, because partnerships that pretend it is free tend to abandon it when the costs arrive. For researchers, the costs are mainly time and control. Joint decision-making is slower. Community review of instruments and manuscripts adds months. A community partner's objection can remove a variable a researcher wanted, delay a paper past a promotion deadline, or end a line of inquiry altogether. Researchers also carry reputational risk inside their institutions, where colleagues may regard participatory work as less rigorous, and where tenure committees may not know how to credit a jointly authored report written for a city council. For community partners the costs are different and often larger. Participation consumes the time of people who are already stretched: staff of small organizations with no slack, residents with jobs and caring responsibilities. Community partners carry the relational risk of the project, because it is their credibility with neighbors that is spent if the research disappoints or harms. They are frequently asked to educate researchers about their community, a form of labor that is seldom paid and rarely acknowledged. And when a project ends, the researchers move to the next grant while community partners remain to answer for what happened. Recognizing these asymmetries is part of sharing power. A partnership that compensates the time of community partners, that protects their credibility by honoring agreements, and that accepts slower timelines as the price of legitimacy has done more to share power than one with an elaborate governance chart and no budget line for community staff. When partners disagree Genuine power-sharing means that the community can say no, and eventually it will. A community partner may object to a question that researchers consider scientifically essential, refuse to allow a finding to be published, or conclude that the partnership is no longer worth its time. How a partnership handles these moments reveals what it actually is. The agreement should anticipate them. Common provisions include a period of discussion before any partner can withdraw, a commitment to seek mediation from a trusted third party, and rules about what happens to data and publications if the partnership dissolves. Some agreements distinguish between the community's right to review and comment on all publications and its right to veto publications that would identify or harm the community; others give a broader veto. Researchers sometimes worry that such provisions threaten academic freedom. The counterargument is that academic freedom protects researchers from interference by their employers and the state; it does not entitle them to publish data that others contributed on condition of shared control. Where the conflict is real, the honest course is to negotiate the terms before the research begins, when both parties can walk away without loss. Disagreement is not failure. Jagosh and colleagues found that partnerships which worked through conflict openly often emerged stronger, because the experience demonstrated that community voices had real weight. What damages partnerships is not disagreement but the discovery that the community's agreement was never actually required. Chapter 3. Whose Question Is It? Of all the decisions in a research project, the choice of question is the one that most constrains every other. It determines what data will be collected and what will be ignored, which people will be counted and which variables will explain them, what kind of answer is possible and to whom it will be useful. A community brought in after the question has been fixed can improve the recruitment materials, suggest better wording for a survey and help interpret the results, but it cannot change what the study is about. For that reason the question is where participatory research most often fails without anyone noticing, because the failure takes the form of an absence: the question the community would have asked and never got to. How questions are usually made In conventional research, questions come from three sources: the researcher's discipline, which defines what is interesting and publishable; the funder, which defines what will be paid for; and the researcher's own career, which rewards questions that extend a line of work. None of these sources is illegitimate. Disciplines accumulate knowledge by building on previous findings; funders have mandates; researchers need to specialize. But none of them runs through the people whose lives the research is about. The consequences are visible in what gets studied. Funding in health research is organized largely by disease and organ system, while people experience their health through housing, work, income, neighborhood safety and the behavior of institutions. A community asked to partner on a diabetes study may care more about the absence of a grocery store, the cost of insulin, or the fact that the clinic closes before people get home from work. Those are researchable questions, but they fall between funding streams, and a researcher with a diabetes grant has limited room to pursue them. Framing also matters within a topic. Research on marginalized communities has a long tradition of what critics call deficit framing: asking what is wrong with the community, its behaviors, its knowledge, its compliance, rather than what is wrong with the conditions and institutions it faces. A study asking why residents fail to attend appointments will look for explanations in residents; a study asking what the clinic does that makes attendance difficult will look in the clinic. The data may overlap, but the findings, and the interventions they justify, will differ. Israel and colleagues' principle that CBPR builds on strengths and resources within the community is partly a corrective to this habit. It asks researchers to begin from what a community already does to protect its health, and from the conditions that undermine those efforts. Communities that arrive with questions Some of the most consequential participatory research began when a community already had a question and could not get anyone to take it seriously. The Flint water crisis is the best-known recent case. After the city of Flint, Michigan, switched its water source to the Flint River in April 2014 under state-appointed emergency management, residents complained of discolored, foul-smelling water, rashes and hair loss. Officials repeatedly assured them the water was safe. One resident, LeeAnne Walters, whose household water tested very high for lead, contacted Marc Edwards, a civil engineer at Virginia Tech who had previously exposed lead contamination in Washington, DC. In 2015 the Virginia Tech team and Flint residents organized a citywide sampling effort in which residents collected water samples from their own homes using kits the researchers supplied, and the results showed lead levels well above federal action levels in many homes. At about the same time the pediatrician Mona Hanna-Attisha and colleagues analyzed children's blood lead records and found that the proportion of children with elevated blood lead levels had increased after the water switch, most sharply in the neighborhoods with the highest water lead. The two lines of evidence together forced officials to acknowledge the problem. Flint is not a textbook CBPR partnership; the collaboration was assembled in an emergency and some residents later criticized how credit and attention were distributed. But it shows with unusual clarity what happens when research follows a community's question. The residents already knew something was wrong. Official monitoring had been designed, and in some respects conducted, in ways that failed to detect it. What the researchers contributed was not the question but the capacity to answer it in a form that institutions could not dismiss. The environmental justice movement has produced many such cases, because environmental hazards are often first detected by the people who live beside them. In California's San Joaquin Valley, residents of small rural communities, many of them Latino farmworker households, had long worried about nitrate and arsenic in their drinking water. Working with the Community Water Center, a local advocacy organization, researchers including Carolina Balazs and Rachel Morello-Frosch at the University of California, Berkeley, examined the distribution of contamination across community water systems and found that systems serving higher proportions of Latino residents and renters tended to have higher nitrate levels. Balazs and Morello-Frosch later used the experience to argue, in a 2013 paper in the journal Environmental Justice, that community participation strengthens what they called the "three Rs" of science: its rigor, relevance and reach. Community partners, they found, directed attention to the right questions: not only whether water exceeded regulatory limits, but who was exposed, who paid for the costs of treatment and replacement water, and which small systems lacked the capacity to fix their problems. Those questions made the research both more accurate about the real burden and more useful for policy. Methods for finding the question together When a partnership begins without a pre-formed question, it needs ways to generate one jointly. Several approaches have become established. Community assessment brings together existing data, from health departments, census records and service providers, with residents' own accounts, gathered through listening sessions, community forums or door-to-door conversations. The important feature is that residents help interpret the data rather than simply supplying stories to illustrate it. A map of asthma hospitalizations means one thing to an epidemiologist and something richer to residents who know which blocks have mold-infested housing, which schools sit next to truck routes and which landlords ignore complaints. Photovoice, developed by Caroline Wang and Mary Ann Burris and described in a 1997 article in Health Education & Behavior, gives participants cameras to document their own community's strengths and concerns, followed by group discussion of the photographs and, typically, an exhibition for policymakers. Wang and Burris drew explicitly on Freire's idea of critical consciousness and on feminist theory. Photovoice has been used thousands of times since, sometimes as a data collection method within a study defined elsewhere, but its original purpose was agenda-setting: letting people who are rarely asked show what they think matters. Priority-setting partnerships formalize the process at a larger scale. The James Lind Alliance, established in the United Kingdom in 2004 and now part of the National Institute for Health and Care Research, brings together patients, carers and clinicians to identify and rank the most important unanswered questions about a particular condition. Its method moves from an open survey of uncertainties, through checking which are genuinely unanswered in the existing literature, to a facilitated workshop at which participants agree a top ten. The resulting lists have repeatedly diverged from the questions researchers had been funding, with patients and carers placing more weight on quality of life, side effects and practical management than on new drugs. Funders, including the National Institute for Health and Care Research itself, have used these lists to commission studies. Concept mapping and deliberative workshops offer other structured ways to reach agreement. What these methods share is a sequence: open generation of concerns by those affected, joint sorting and prioritization, and a deliberate step in which researchers help translate priorities into questions that can be answered with available methods and resources, while the community checks that the translation has not changed the meaning. Translation without capture That last step is where control is most easily lost. Residents may name a concern such as "our kids can't breathe here," which must be turned into a research question: what exposures, measured how, compared with what, over what period, linked to what outcomes? Every one of these choices involves technical judgment, and the person making it holds power. If researchers make the choices alone, the resulting question may be answerable but no longer the community's. Good practice keeps the translation visible. Researchers lay out options and their consequences in plain terms: measuring particulate matter at fixed monitors is cheaper and comparable with official data, while personal monitors carried by children capture actual exposure but are costlier and burdensome; comparing hospitalization rates across neighborhoods is fast but will not show what happens inside homes. The community chooses among options with an understanding of the trade-offs, and it can insist on the more expensive or slower design if that is the one that answers its question. Where no feasible design can answer the community's question, researchers say so, and the partnership decides whether a narrower study is still worth doing. Translation also involves outcomes. What would count as an answer, and what would be done with it? Communities often want research that can be used in a specific venue, a zoning hearing, a school board decision or a funding application, and the design must produce evidence in a form that venue will accept. This is not a corruption of science. It is the same consideration that leads a pharmaceutical trial to use endpoints a regulator will recognize. Questions that harm Communities sometimes resist questions not because they lack interest in the topic but because they have learned what such questions can do. Research on stigmatized conditions, such as alcohol use, mental illness, HIV or violence, produces findings that attach to the whole group studied, including people who never took part. A community-level finding travels further than any individual's data, and it can be used by outsiders to justify prejudice, disinvestment or intervention. The Barrow Alcohol Study is a standard warning. In 1979 researchers from the University of Pennsylvania, working under contract to the North Slope Borough in Alaska, surveyed residents of the Iñupiat town of Barrow (now Utqiaġvik) about drinking. The results were released at a press conference in 1980 before the community had reviewed them, and national newspapers reported that alcoholism was threatening the survival of the Iñupiat people. The borough's bond rating reportedly suffered, residents felt publicly humiliated, and the episode became a case study in research ethics courses. Edward Foulks, one of the investigators, later wrote an account of what went wrong, describing a partnership in which the community had commissioned the research but had no say over how its findings were framed and released. The data might have been accurate; the harm lay in a framing that described the community through its pathology and in a release process that bypassed the people described. The lesson for question-setting is that communities have a legitimate interest in how they will be represented, and that this interest is engaged at the start, not only at publication. A community may agree to study alcohol use if the question is framed around the availability of treatment, the effect of alcohol outlet density or the success of community-led sobriety movements, and refuse if it is framed as the prevalence of drinking. Such preferences can look to researchers like spin. Often they are an accurate recognition that the choice of what to measure determines what story will be told, and that the community, not the researcher, will live with the story. Group harm is poorly handled by conventional research ethics, which focus on the risks to individual participants. An individual can consent to answer questions about her drinking; she cannot consent on behalf of her town to be described as a town of drinkers. Participatory research offers one of the few mechanisms through which a group can weigh these risks collectively, and it does so best at the moment the question is chosen. Questions researchers would rather not ask The reverse problem also arises. Communities sometimes want questions that researchers, or their institutions, would rather avoid. Residents near a university may want to study the university's own effects on housing costs and displacement. Patients may want to examine how a hospital treats them. Communities policed heavily may want research on police conduct rather than on crime. Workers may want to study their employer's safety practices, and the employer may be a research funder. These questions test whether a partnership's commitment to community priorities holds when the priorities become inconvenient. The institution may have legitimate concerns about conflicts of interest, and researchers may fear retaliation from colleagues or funders. But a partnership that systematically steers away from questions implicating the institutions involved has a structural bias worth naming. One sign of a mature partnership is that it has at least once pursued a question that made its academic partner uncomfortable, and survived. When the funder has already chosen A frequent difficulty is that the question is partly fixed before any partnership exists. Funding announcements define topics, and applications have deadlines that leave little time for community deliberation. Researchers then approach community organizations with a grant already half-written and ask them to sign a letter of support. There are honest ways to handle this. The first is to build partnerships before funding opportunities arise, so that when an announcement appears, the partnership already has a list of priorities and can decide together whether this opportunity matches any of them. Long-standing partnerships such as the Detroit center operated in exactly this way, reviewing proposed projects against the partnership's agreed priorities. The second is to be explicit about what is fixed and what is open: the topic is set by the funder, but within it the community will choose the specific question, the population and the outcomes. The third is to decline. A community organization that turns down an opportunity because it does not serve its priorities is exercising exactly the power participatory research claims to respect, and researchers who treat such refusals as obstacles have misunderstood the enterprise. Some funders have changed their practices in response. Planning grants that pay for partnership development before a full proposal, application processes that require evidence of community involvement in defining the question, and review panels that include community members all shift the point at which communities enter. The Patient-Centered Outcomes Research Institute in the United States, for example, has required applicants to describe how patients and other stakeholders were involved in developing the research question, and includes patients and stakeholders as reviewers alongside scientists. Such measures do not guarantee genuine partnership, but they make it harder to recruit a community after the fact. The question is the first allocation of power in a research project. Once it is made, the choice of data follows, and with it a new set of decisions about who collects that data, on what terms, and at whose risk. That is the subject of the next chapter. Hashtags: #CommunityBasedParticipatoryResearch #CBPR #ParticipatoryResearch #CommunityEngagement #CollaborativeResearch #ParticipatoryActionResearch #ActionResearch #CommunityAcademicPartnerships #PowerSharing #SharedDecisionMaking #CommunityGovernance #ResearchEquity #SocialImpact #CommunityEmpowerment #DecolonizingMethodologies #IndigenousResearch #Photovoice #CommunityAssessment #PrioritySetting #ResearchPartnerships #DataOwnership #CommunityAuthorship #CoLearning #HealthEquity #FutureOfParticipatoryResearch
- Time Series Analysis for Physical Sciences (Wavelets, ARIMA, and Spectral Decomposition)
Download the Book (PDF): Introduction In 1927 the Cambridge statistician George Udny Yule sat down with roughly a century and a half of annual sunspot numbers and asked a question that still defines the subject of this book. Astronomers had known since Heinrich Schwabe's observations in the 1840s that sunspots wax and wane on a cycle of roughly eleven years, and the natural instinct was to describe the record as a sum of periodic components, a few sine waves of fixed period, amplitude and phase, with some observational error sprinkled on top. Yule thought that picture was wrong. The cycle drifted. Some cycles lasted nine years, others fourteen. Peaks came in different heights for no reason that any fixed set of sines could capture. He proposed instead that the sunspot record behaved like a pendulum being pelted at random by boys with peashooters: a system with its own natural tendency to oscillate, continually knocked about by disturbances, so that its rhythm persisted but never locked into strict periodicity. From that image he built what we now call an autoregressive model, in which each year's value is a weighted combination of the previous years' values plus a random shock. Yule's paper is usually remembered as the birth of autoregressive modelling. It deserves to be remembered for something more basic. It was an argument that the method one uses to analyse a time series is not a neutral lens. Describing the sunspot record as a sum of sines amounts to claiming that the Sun contains clocks. Describing it as an autoregressive process amounts to claiming it contains a damped oscillator driven by noise. The two descriptions can fit the same data tolerably well, yet they imply different physics, different predictions, and different answers to the question of whether a particular peak in the record is meaningful. That is the idea this book is organised around. Every technique for decomposing a time series carries within it a model of how the series was generated, and in the physical sciences, where records are noisy, reddened, non-stationary and often irregularly sampled, the analyst's first responsibility is to make that model explicit, match it to the physics, and test every apparent signal against a credible model of the noise. A spectral peak, a patch of high wavelet power, a significant coherence, or a fitted trend means nothing on its own. It becomes evidence only once it has beaten an honest null hypothesis, and in geophysics, oceanography and astrophysics that null is almost never white noise. Why physical time series are hard Physical scientists come to time-series analysis from a different direction than economists or engineers. An electrical engineer designing a filter usually controls the signal and can repeat the experiment. An economist forecasting quarterly output usually cares more about the next few values than about the mechanism that produced them. A physical scientist typically has one record of a natural process that no one controls, cannot be repeated, and was sampled by instruments that changed over time. The goal is rarely forecasting for its own sake. It is inference: is there a periodicity here, how strong is it, does it change, is it linked to that other record, and what does that say about the underlying system? Several features of physical records conspire to make this difficult. The first is redness. Most geophysical and many astrophysical processes have memory. Ocean temperature anomalies persist because the mixed layer has large heat capacity. Klaus Hasselmann showed in 1976 that a slow system integrating fast random weather forcing will naturally produce a spectrum whose power rises toward low frequencies, with no oscillatory mechanism required. Accreting black holes, stellar granulation and the geomagnetic field all show similar red or "flicker" behaviour. Against a red background, random fluctuations at low frequencies are large, and a naive test that assumes white noise will declare spurious long-period cycles significant with depressing regularity. The second is non-stationarity. The statistical properties of the process change over time. The amplitude of the El Niño–Southern Oscillation varies from decade to decade. Solar cycles differ in length and strength. A seismometer records hours of ambient microseismic noise and then a few minutes of an earthquake whose frequency content evolves as different wave phases arrive. Methods built on the assumption that the same spectrum applies throughout the record, which includes the classical Fourier periodogram and standard ARMA models, can average over precisely the changes one wants to detect. The third is irregular and imperfect sampling. Ground-based astronomical observations are interrupted by daylight, weather and the phases of the Moon. Sediment and ice cores yield samples that are evenly spaced in depth but not in time, and the conversion from depth to age is itself uncertain. Tide gauges fail, moorings are serviced, satellites change orbit. Many elegant methods assume regular sampling with no gaps, and the ways they fail when that assumption is violated are rarely obvious from their output. The fourth is the mixing of timescales. A single record may contain a secular trend, a seasonal cycle, interannual oscillations, and high-frequency noise, and the analyst usually wants to isolate one of them without distorting the others. How one removes a trend affects the spectrum at low frequencies; how one handles the seasonal cycle affects everything near a year; how one filters affects the apparent phase relationships between records. Separating timescales is never a purely technical preprocessing step. It embeds assumptions about what the components are. What this book covers The chapters move from the foundations to the more specialised tools, and each one is organised around what the method assumes and how it fails when those assumptions do not hold. Chapter 1 sets up the vocabulary: sampling, the Nyquist limit and aliasing, stationarity in its various senses, autocorrelation, and the families of noise that physical records display. It makes the case that choosing a noise model is the most consequential decision in most analyses. Chapter 2 develops Fourier analysis from the ground up, explaining what the discrete Fourier transform actually computes, why the Fast Fourier Transform made spectral analysis practical, and what the finite length of any real record does to the result. Leakage, windowing and resolution are treated as consequences of a single fact rather than as a list of rules. Chapter 3 turns the Fourier transform into a statistical tool. The raw periodogram is a poor estimator of the spectrum, and the chapter explains the main remedies, including smoothing, Welch's segment averaging, and Thomson's multitaper method, before tackling the question that matters most in practice: how to decide whether a peak is real when the background is red. Chapter 4 returns to Yule's idea. It develops autoregressive, moving-average and integrated models, explains how they are identified and fitted, shows how an autoregressive model defines a spectrum of its own, and discusses unit roots and differencing, which are far more consequential for physical records than their origins in econometrics might suggest. Chapter 5 introduces the continuous wavelet transform as the natural response to non-stationarity. It explains the trade-off between time and frequency resolution, the choice of mother wavelet, the cone of influence, and the significance testing framework that Christopher Torrence and Gilbert Compo made standard in the geosciences. Chapter 6 extends wavelets to pairs of records: cross-wavelet power, wavelet coherence and phase. These tools are widely used and widely misread, and the chapter is as much about what coherence cannot tell you as about what it can. Chapter 7 deals with trends, from ordinary regression with autocorrelated residuals through smoothing filters, the Hodrick–Prescott filter, and the sparse methods known as ℓ1 trend filtering, to data-adaptive decompositions such as singular spectrum analysis and empirical mode decomposition. Its central argument is that a trend is a modelling choice about timescales, not an object waiting in the data to be found. Chapter 8 addresses irregular sampling, gaps and uncertain time axes, which are the everyday reality in observational astronomy and paleoclimatology. It covers the Lomb–Scargle periodogram and its pitfalls, continuous-time autoregressive models and Gaussian processes, and the propagation of chronological uncertainty into spectral conclusions. The examples are drawn throughout from the three fields named in the subtitle and their neighbours: tides and ocean temperature, earthquakes and the Earth's rotation, variable stars and quasars, ice cores and the ice ages. They are chosen because they show a method working, or failing, for a reason that generalises. How to read it The book assumes familiarity with basic calculus, complex numbers, and elementary probability, and some exposure to linear regression. It does not assume prior study of signal processing. Equations are kept to the minimum needed to make an idea precise, and where one appears, the text explains what each term means physically. It is not a software manual. Excellent implementations of every method discussed here exist in Python, R, MATLAB and Julia, and they change faster than any book. What changes more slowly is the judgement required to use them: knowing which method's assumptions fit the problem, what the output is actually estimating, and how to recognise when a plausible-looking result is an artefact. That judgement is what this book tries to supply. A final note on attitude. Time-series analysis in the physical sciences has a long history of false discoveries: cycles in sunspots that predicted wheat prices, periodicities in earthquake catalogues, spectral peaks in paleoclimate records that dissolved when the age model was revised. Most of these were not the result of carelessness. They came from using methods correctly under assumptions that did not hold. The discipline that prevents them is not scepticism for its own sake but a habit of asking, at every step, what the method is assuming about the process that made the data, and whether nature agrees. Chapter 1. Records of a Restless World A time series is a sequence of measurements ordered in time, but that bland definition hides most of what matters. The order is not incidental. In a sample of independent measurements, such as the heights of a thousand people, shuffling the values changes nothing. In a time series, shuffling destroys the very information one is trying to extract, because the values are related to each other through the dynamics of the system that produced them. Today's sea-surface temperature is close to yesterday's because the ocean has heat capacity. The brightness of a pulsating star at one moment constrains its brightness a fraction of a cycle later because the star is a resonant body. The whole enterprise of time-series analysis consists of describing and exploiting that dependence. This chapter sets out the basic vocabulary that the rest of the book relies on: how continuous processes become discrete records, what stationarity means and why it is nearly always violated, how dependence is measured, and what kinds of noise physical systems produce. The argument running underneath is that before any decomposition is attempted, the analyst needs a working hypothesis about what the record would look like if it contained nothing interesting at all. From continuous process to discrete record Almost every physical process evolves continuously, but we observe it at discrete instants. A tide gauge might record water level every six minutes; a broadband seismometer samples ground velocity at 100 or 200 times per second; a satellite photometer such as that on NASA's Kepler mission integrated stellar brightness over roughly thirty-minute intervals for its long-cadence data and about one minute for short cadence; an ice core might yield one isotope measurement per centimetre of depth, which near the bottom of a deep core can represent centuries. The sampling interval, usually written Δt, sets a hard limit on what can be learned. The highest frequency that can be represented unambiguously is the Nyquist frequency, equal to one over twice the sampling interval. With hourly data, the Nyquist frequency is one cycle per two hours; nothing that oscillates faster can be resolved. This result, associated with Harry Nyquist's work on telegraph transmission in 1928 and Claude Shannon's sampling theorem of 1949, is not a limitation of any particular method. It follows from the fact that a sine wave sampled at discrete points is indistinguishable from infinitely many other sine waves passing through the same points. What happens to variability above the Nyquist frequency is the more insidious problem. It does not vanish. It is aliased: folded back into the resolvable band and disguised as a lower frequency. The classic geophysical example involves satellite altimetry. The TOPEX/Poseidon mission, launched in 1992, revisited each point on its ground track every 9.9156 days. The principal lunar semidiurnal tide, M2, has a period of about 12.42 hours, far too short to be resolved by a ten-day repeat. Sampled at that interval, the M2 tide appears as a spurious oscillation with a period of about 62 days. The mission designers chose the orbit deliberately so that the aliased periods of the major tidal constituents would be distinct from each other and from the annual cycle, allowing the tides to be estimated from the altimeter record. Aliasing was not avoided; it was engineered to be harmless. In most situations there is no such luxury, and the only defence is to filter the continuous signal before sampling, which instrument designers do with analogue anti-aliasing filters. When data are later decimated, for example when minute-resolution records are reduced to hourly averages, the same rule applies: averaging over the interval is a crude low-pass filter, while simply picking every sixtieth value is not a filter at all and will alias whatever high-frequency variability the record contains. The other limit set by sampling is at the low-frequency end. A record of length T cannot resolve periods much longer than T, and it cannot distinguish two frequencies separated by less than about 1/T. The second fact is known in tidal analysis as the Rayleigh criterion. To separate the M2 tide from the principal solar semidiurnal tide, S2, whose period is exactly 12 hours, requires a record at least as long as the beat period between them, about 14.8 days, the spring–neap cycle. To separate the diurnal constituents K1 and P1 requires about half a year. Oceanographers learned long ago that a month of tide-gauge data cannot yield all the constituents that a year of data can, however clever the analysis. Stationarity and its discontents Most of classical time-series theory rests on the assumption of stationarity. A process is strictly stationary if its joint statistical properties are unchanged by a shift in time: the probability distribution of the values at any set of times is the same as at those times shifted by any fixed lag. A weaker and more practical notion, second-order or weak stationarity, requires only that the mean is constant and that the covariance between values depends only on the lag separating them, not on absolute time. Weak stationarity matters because it is exactly what is needed for the power spectrum to be well defined and unchanging. If the covariance structure depends only on lag, then the process can be described by a single function of frequency, its spectral density, which says how the variance is distributed among timescales. This is the Wiener–Khinchin relationship, which states that the spectral density and the autocovariance function are Fourier transforms of each other. It is the theoretical backbone of Chapters 2 and 3. Physical records violate stationarity in many ways, and it helps to distinguish them, because they call for different remedies. A deterministic trend is a systematic change in the mean, such as the rise of global mean sea level or the growth of atmospheric carbon dioxide measured at Mauna Loa since Charles David Keeling began the record in 1958. Removing it with a fitted function can leave a residual that is approximately stationary. A stochastic trend, or unit root, is a subtler kind of wandering in which shocks accumulate permanently rather than decaying. A random walk is the simplest example. Its variance grows without bound over time, and no fitted curve will make it stationary; differencing is required. Chapter 4 explains why distinguishing these two cases matters and why it is often difficult. A periodic modulation of statistics, sometimes called cyclostationarity, occurs when the variance or correlation structure itself varies with a known cycle. Weather variability is larger in some seasons than others; the amplitude of diurnal temperature swings depends on cloud cover, which depends on season. Removing the mean seasonal cycle does not remove this. Evolving spectral content is the case that wavelets are designed to handle. The frequency of an oscillation drifts, or its amplitude comes and goes. The El Niño–Southern Oscillation was notably more active in some decades of the twentieth century than others. Solar cycles vary in length between roughly nine and fourteen years. A seismogram is non-stationary almost by definition. Finally, there are abrupt changes: instrument replacements, changes in observing practice, station relocations, volcanic eruptions. Homogenisation of climate records, the detection and correction of such breakpoints, is a discipline of its own, but the analyst of any long record should assume that such breaks exist until shown otherwise. It is tempting to treat stationarity as a box to tick before proceeding. A better attitude is to treat it as a modelling choice made at a particular timescale. A record may be reasonably stationary over a few years and clearly non-stationary over a century. The question is not whether the process is stationary in some absolute sense, but whether the assumption is adequate for the question being asked over the span of data available. Measuring dependence The basic tool for describing how a stationary series depends on its own past is the autocorrelation function, which gives the correlation between values separated by each lag. For white noise, a sequence of independent values, the autocorrelation is one at lag zero and zero at every other lag. For a record with memory, the autocorrelation decays gradually, and the rate of decay characterises the memory. The sample autocorrelation, computed from data, is a noisy estimate, and its noise is itself correlated across lags. This is one reason why reading structure into the wiggles of a sample autocorrelation function is hazardous. A useful rule is that for a white-noise process of length N, sample autocorrelations at nonzero lags are roughly normal with standard deviation one over the square root of N, so values beyond about twice that are notable. But the rule applies only to white noise. When the series has memory, the sampling variability of the autocorrelation at longer lags is larger, and apparent oscillations in the autocorrelation function of a red-noise process are common. The practical consequence of dependence that every physical scientist should internalise concerns effective sample size. If a record of N values has positive autocorrelation, those N values contain less independent information than N independent measurements would. For a first-order autoregressive process with lag-one autocorrelation r, a widely used approximation gives an effective sample size for estimating the mean of roughly N times (1 − r) divided by (1 + r). With a lag-one autocorrelation of 0.8, typical of monthly ocean temperature anomalies, 600 monthly values carry about as much information about the mean as 67 independent values. Benjamin Santer and colleagues, in a 2000 paper in the Journal of Geophysical Research, applied this kind of adjustment to the uncertainty of temperature trends and showed how much it widened confidence intervals. Ignoring autocorrelation makes every significance test in the analysis overconfident, and the overconfidence can be severe. Cross-dependence between two series is measured by the cross-correlation function, the correlation between one series and the other shifted by each lag. The same warning applies with greater force. Two independent red-noise series will frequently show large sample cross-correlations purely by chance, because each has few effective degrees of freedom. Yule himself identified the problem in a 1926 paper on "nonsense correlations" between time series, and it has been rediscovered in every generation since. The colours of noise If one had to choose a single concept to carry away from this chapter, it would be that noise in physical systems has a spectrum, and the spectrum is rarely flat. White noise has equal power at all frequencies, named by analogy with white light. It is the noise of independent measurement errors and, approximately, of many instrumental processes at high frequencies. Its autocorrelation vanishes at all nonzero lags. Red noise has power increasing toward low frequencies. The term is used in two related senses. In the loose sense it means any spectrum that rises toward low frequencies. In the specific sense favoured in climate science, it means the spectrum of a first-order autoregressive process, AR(1), in which each value equals a fraction of the previous value plus a white-noise shock. The AR(1) spectrum is flat at the lowest frequencies, falls off at higher frequencies, and has a characteristic "knee" at a frequency set by the memory timescale. It is the discrete-time analogue of a continuous process known to physicists as the Ornstein–Uhlenbeck process, the velocity of a Brownian particle subject to friction. Hasselmann's 1976 insight was that the slow components of the climate system integrate fast weather noise in exactly this way, producing an AR(1)-like spectrum without any internal oscillation. That result turned red noise from a nuisance into a physically motivated null hypothesis. Power-law noise has a spectrum proportional to frequency raised to a negative power. When the exponent is two, the process is a random walk, sometimes called Brownian or brown noise. When it is one, the process is called flicker noise or pink noise; it appears in electronic devices, in the brightness fluctuations of stars and quasars, in some sea-level and geodetic records, and in many other systems. Power-law processes lack a single characteristic timescale, and their correlations decay slowly. Those with exponents between zero and one are long-memory processes, a phenomenon first noticed in hydrology when Harold Edwin Hurst, studying Nile flood records for the design of reservoirs, found in 1951 that the range of cumulative departures grew faster with record length than independent or short-memory processes would allow. The spectral shape of noise matters in practice because it sets the background against which any signal must be judged. A peak at a period of twenty years in a century-long record looks impressive if one compares it with a flat spectrum. Compared with a red-noise spectrum fitted to the same data, it may be entirely ordinary. The history of climate-cycle claims is full of peaks that failed this comparison once it was made properly. Michael Mann and Jonathan Lees, in a 1996 paper in Climatic Change, proposed methods for estimating a robust red-noise background precisely because so many claimed periodicities in climate records dissolved against it. The noise model also affects estimated uncertainties of trends and other parameters. In geodesy, analyses of continuous GPS position time series in the late 1990s and 2000s found that the noise contained a substantial flicker component, and that assuming white noise understated the uncertainty of station velocities by large factors. The same lesson has been learned in sea-level analysis and in the estimation of warming trends. A record read three ways The Mauna Loa carbon dioxide record shows how these ideas interact in a single, familiar series. Measured continuously since 1958, the monthly means display three obvious components: a rising mean, a regular seasonal cycle with an amplitude of a few parts per million, and small irregular fluctuations. Each component raises one of the issues above. The rise is non-stationarity of the most obvious kind, and it is not linear: the growth rate has increased over the decades, from under one part per million per year in the early record to more than two in recent years. Fitting a straight line would leave a curved residual that masquerades as low-frequency variability in any subsequent spectrum. The seasonal cycle, produced mainly by the uptake and release of carbon by Northern Hemisphere vegetation, is close to periodic but not exactly so; its amplitude has grown slowly over the record, a change that has been linked to shifts in northern ecosystems. Removing a fixed mean seasonal cycle therefore leaves a residual annual signal of slowly changing size. The irregular fluctuations in the growth rate, finally, are not white. They are coherent with the El Niño–Southern Oscillation, because the terrestrial biosphere takes up less carbon in the warm, dry conditions that El Niño brings to much of the tropics, and they show the persistence typical of a system with memory. Every decision about how to separate these components is a decision about the process. Treating the rise as a smooth deterministic curve, the seasonal cycle as a slowly evolving harmonic, and the remainder as red noise is a physically defensible choice, but it is a choice, and the variability attributed to each component depends on it. The same will be true of every record in this book. Choosing the null before looking for the signal It follows from all this that a physical time-series analysis should begin by stating its null hypothesis explicitly. The null is the model of what the data would look like if the phenomenon of interest were absent. It should be physically plausible for the system in question, estimated from the data or from independent information, and simple enough to be tested against. For climate and ocean records the default null is usually AR(1) red noise, occasionally something with more structure. For astronomical light curves it may be white photon noise plus a red component from stellar granulation or accretion variability, often described by a damped random walk or a sum of damped oscillators. For seismic and geodetic records it may be a mixture of white and flicker noise. For sunspots and similar records it may be a noise-driven damped oscillator of exactly the kind Yule proposed. Once the null is chosen, the whole analysis can be framed as a comparison. A spectral peak is significant if it would rarely arise under the null. A wavelet power feature is significant if its magnitude exceeds what the null would produce at that scale. A coherence between two records is significant if it exceeds what two independent realisations of the null would show. A trend is significant if it is unlikely under a null with the same noise properties but no trend. This framing has a cost. It forces the analyst to confront the possibility that the most interesting feature in the data is noise, and it often yields less exciting results than a naive analysis would. It also has limits: a null that is too flexible can absorb real signals, and choosing among several plausible nulls introduces its own subjectivity. But it is the only framing that gives spectral, wavelet and trend results a clear meaning, and the chapters that follow return to it repeatedly. Chapter 2. From Sums of Sines to the Fast Fourier Transform The idea that a complicated function can be built from simple oscillations is older than the calculus used to prove it. Joseph Fourier's work on heat conduction, presented to the French Academy in 1807 and published in 1822 as Théorie analytique de la chaleur, argued that essentially any function on an interval could be represented as a sum of sines and cosines. Mathematicians spent the rest of the nineteenth century working out precisely when that claim is true, but the physical insight was sound: many linear systems respond to each frequency independently, so decomposing a signal by frequency often decomposes the physics. For time-series analysis the Fourier idea offers something more specific. It gives a way to ask how the variance of a record is distributed among timescales. A record with a strong annual cycle will have much of its variance at one cycle per year; a record dominated by slowly varying red noise will have most of its variance at low frequencies; a record of white measurement error will have its variance spread evenly. This chapter explains what the discrete Fourier transform computes, how the Fast Fourier Transform made it practical, and what the finite length of a real record does to the answer. What the discrete Fourier transform computes Suppose we have N equally spaced values of a series. The discrete Fourier transform re-expresses these N numbers as the amplitudes and phases of sinusoids at N/2 + 1 specific frequencies, running from zero up to the Nyquist frequency in steps of 1/(NΔt). These are the Fourier frequencies. The lowest nonzero one completes exactly one cycle over the record; the next completes two; and so on up to the Nyquist frequency, at which the sinusoid alternates sign from one sample to the next. Several properties of this transformation are worth stating plainly, because they explain most of what follows. First, the transformation is exact and invertible. No information is lost. The N data values and the N real numbers describing the Fourier coefficients (a real part and imaginary part at each frequency, with the zero and Nyquist frequencies having only real parts) are two descriptions of the same thing. Second, the Fourier basis functions are orthogonal over the record. This means that computing the coefficient at one Fourier frequency is equivalent to a least-squares fit of a sine and cosine at that frequency, and that fit is unaffected by what is happening at the other Fourier frequencies. It also means that the total variance of the record is exactly the sum of the variances attributed to each frequency, a result known as Parseval's theorem. This is what licenses reading the squared Fourier amplitudes as a decomposition of variance. Third, the transform implicitly treats the record as one period of an infinitely repeating signal. The Fourier frequencies are exactly the frequencies that fit a whole number of cycles into the record length, so any sum of them is periodic with period equal to the record length. If the last value of the record is very different from the first, the implied periodic extension has a jump at the join, and that jump must be represented by the Fourier components. This single fact, that the DFT sees a finite record as one cycle of a periodic signal, underlies leakage, the need for windowing, and several other effects discussed below. The squared magnitudes of the Fourier coefficients, suitably scaled, form the periodogram, first used under that name by the physicist Arthur Schuster in 1898 in a search for hidden periodicities in meteorological and later sunspot data. The periodogram is the most direct estimate of how the variance is distributed across frequency. Chapter 3 explains why it is also a surprisingly poor one. The Fast Fourier Transform Computed directly from its definition, the DFT of N points requires on the order of N squared complex multiplications: N coefficients, each a sum over N values. For a record of a million points, typical of a day of high-rate seismic data or a long satellite light curve, that is a trillion operations. For most of the twentieth century this cost confined spectral analysis to short records or indirect methods. In 1965 James Cooley of IBM and John Tukey of Princeton published "An algorithm for the machine calculation of complex Fourier series" in Mathematics of Computation. They showed that when N is a power of two, the DFT can be split into two DFTs of half the length, one on the even-indexed values and one on the odd-indexed values, which are then combined with a small number of extra multiplications. Applying the split recursively reduces the cost to the order of N times the logarithm of N. For a million points the saving is a factor of about fifty thousand. The history has a physical-science twist. The motivation that reached Cooley came partly from the need to detect Soviet nuclear tests with seismometers, a problem that required spectral analysis of long seismic records. And the underlying idea turned out to be much older. Michael Heideman, Don Johnson and Sidney Burrus showed in a 1984 article that Carl Friedrich Gauss had used essentially the same decomposition around 1805 to interpolate the orbits of the asteroids Pallas and Juno, in work published only posthumously. Several twentieth-century authors had also rediscovered forms of it. What Cooley and Tukey provided, at a moment when digital computers could exploit it, was a clear general algorithm that launched modern digital signal processing. Today the practical details are handled by mature libraries. FFTW, developed by Matteo Frigo and Steven Johnson at MIT and described in a 2005 paper in the Proceedings of the IEEE, adapts its algorithm to the length of the data and the hardware it runs on and handles lengths that are not powers of two efficiently. Padding a record with zeros to reach a power of two is therefore no longer necessary for speed. It remains common for other reasons, and as the next section explains, it does not do what many users believe it does. Leakage, resolution and zero padding Consider a record containing a single pure sinusoid whose frequency falls exactly on one of the Fourier frequencies. Its periodogram is zero everywhere except at that frequency. Now shift the sinusoid's frequency so that it falls halfway between two Fourier frequencies. The record no longer contains a whole number of cycles, so its periodic extension has a discontinuity at the join. The periodogram now shows power not only at the two neighbouring Fourier frequencies but spread across the whole spectrum, decaying slowly with distance from the true frequency. This spreading is called spectral leakage. Leakage has a clean mathematical explanation. A finite record is an infinite signal multiplied by a rectangular window that is one during the observation and zero outside it. Multiplication in time corresponds to convolution in frequency, so the spectrum we compute is the true spectrum convolved with the Fourier transform of the rectangular window. That transform, a function of the shape sin(x)/x, known as the Dirichlet or sinc kernel in its discrete form, has a main lobe of width about 1/T and an endless series of sidelobes whose peaks decay only in proportion to the inverse of the distance from the centre. The highest sidelobe is only about 13 decibels below the main peak, roughly a twentieth of its power. Two consequences follow, and they pull in different directions. The width of the main lobe sets the frequency resolution: two sinusoids closer in frequency than about 1/T blur into one. This is the Rayleigh criterion of Chapter 1 in another guise. The height of the sidelobes sets the dynamic range: a weak signal near a strong one can be buried under the strong signal's leakage. In physical records this second problem is usually more serious than it sounds. Geophysical spectra often span many orders of magnitude, with enormous power at low frequencies and little at high frequencies. Leakage from the low-frequency peak can raise the apparent level of the high-frequency spectrum by large factors, flattening its slope and hiding real features. Zero padding means appending zeros to the record before transforming it. It evaluates the spectrum at more closely spaced frequencies, which makes plots smoother and can help locate the frequency of an isolated peak more precisely. It does not improve resolution. The main lobe still has width 1/T, because the observation still lasted only T. Two peaks that were blurred together remain blurred together, however finely the blurred result is sampled. It is a common and understandable confusion, and it leads to overstated claims about separating nearby periodicities. Windows and tapers The standard response to leakage is to multiply the record by a smooth window, also called a taper, that falls to zero at the ends before transforming. This removes the discontinuity at the join and dramatically reduces sidelobes. The price is a wider main lobe, and hence coarser resolution, and a loss of effective data, because values near the ends are down-weighted. The window choice therefore embodies a trade-off between resolution and leakage, and the right choice depends on the spectrum being analysed. Fredric Harris's 1978 survey in the Proceedings of the IEEE catalogued dozens of windows and their properties and remains a standard reference. For most physical-science purposes a handful suffice, and their main properties can be compared on a few criteria, as Table 1 sets out. Table 1. Common data windows and their trade-offs (values from Harris, 1978). Window Highest sidelobe (dB) Sidelobe decay (dB per octave) 3-dB bandwidth (relative to rectangular) Typical use Rectangular (none) −13 −6 1.0 Periodic signals that fit the record exactly; transient analysis Hann −32 −18 about 1.6 General-purpose default Hamming −43 −6 about 1.5 Moderate dynamic range, close-in sidelobes suppressed Blackman–Harris (4-term) −92 −6 about 2.1 Very high dynamic range, steep red spectra The Hann window, a single cycle of a raised cosine, is a sensible default because its sidelobes are low and fall off quickly. The Hamming window suppresses the nearest sidelobes more, but its far sidelobes decay slowly, which matters when the spectrum is steep. When the spectrum spans many orders of magnitude, as with ocean wave spectra or the red spectra of many geophysical variables, a window with very low sidelobes is often worth the loss of resolution. The tapered cosine or Tukey window, which is flat in the middle and tapers only a fraction of the record at each end, offers a middle path, and tapering 10 per cent of the record at each end was long a standard recommendation in geophysical practice. An alternative approach to steep spectra is prewhitening. One first applies a simple filter that flattens the spectrum, most often a first difference or an AR filter fitted to the data, then computes the spectrum of the filtered series with reduced leakage, and finally divides by the known response of the filter to recover the original spectrum. John Tukey and Ralph Blackman recommended this in their 1958 monograph The Measurement of Power Spectra, and it remains effective. It is also an early example of a theme that recurs throughout this book: combining a parametric model of the gross shape of the spectrum with a non-parametric estimate of its details. Chapter 3 introduces a more principled solution to the window problem, Thomson's multitaper method, which uses several optimally designed orthogonal tapers rather than one. Removing the mean and the trend Two preprocessing steps are so routine that they are easily performed without thought. The first is removing the mean. If the record has a nonzero mean, the zero-frequency coefficient will be large, and with any window except the rectangular one, some of that power will leak into the lowest nonzero frequencies. Subtracting the mean before windowing avoids this. The second is removing a linear or polynomial trend. The reasoning is the same: a trend implies a large discontinuity between the end and the start of the periodic extension, which leaks power across the whole spectrum. Detrending before windowing reduces this. But detrending is not innocent. A linear fit removes variance at the lowest frequencies whether or not that variance reflects a genuine trend, and if the process is red noise with no deterministic trend, detrending will artificially depress the lowest frequencies of the spectrum. Chapter 7 returns to this issue in detail. For now it suffices to note that detrending is a statement about the process, not merely a numerical convenience, and it should be reported alongside the spectrum. Linear filtering in the frequency domain The Fourier transform is also the natural language for filtering. A linear, time-invariant filter, such as a running mean, a smoother, or the response of an instrument, acts on each frequency independently: it multiplies the amplitude by some factor and shifts the phase by some amount. The function giving these factors across frequency is the filter's frequency response, or transfer function. Thinking in these terms clarifies many common operations. A running mean over M points has a frequency response of the sin(x)/x shape; it attenuates high frequencies, but not cleanly, and it passes some high-frequency variability with its sign reversed in its negative sidelobes. This is why applying a running mean to noise can create spurious oscillations, a phenomenon identified by Eugen Slutsky in a 1927 paper that, in the same year as Yule's sunspot work, showed that summing random shocks can produce cycle-like fluctuations. Slutsky's result was a warning against reading cycles into smoothed series that remains relevant to anyone who smooths data before looking at it. A filter can also shift phase. A running mean that averages the current and previous M − 1 values introduces a delay of half its length; one centred on the current value introduces none. Recursive filters, which compute each output from previous outputs as well as inputs, are computationally efficient but generally shift different frequencies by different amounts. Running such a filter forward and then backward cancels the phase shift, a technique known as zero-phase filtering, which is appropriate when the whole record is available and the aim is analysis rather than real-time processing. When comparing the timing of events in filtered records, as in Chapter 6, uncorrected filter phase shifts are a frequent source of spurious lags. The frequency-domain view also explains instrument correction. A seismometer's output is ground motion filtered by the instrument's response, which is known from calibration. Dividing the Fourier transform of the recorded signal by the instrument response, with care to avoid amplifying noise at frequencies where the response is small, recovers an estimate of the true ground motion. The same idea underlies deconvolution in many other settings, and its fragility near the edges of an instrument's passband is a practical consequence of the fact that information attenuated to near zero cannot be recovered by division. What the Fourier view assumes It is worth closing by making explicit what the Fourier representation assumes about the process that generated the data, because that sets up the rest of the book. The periodogram of a finite record summarises how the variance of that record was distributed across frequency, averaged over the whole record. It says nothing about when during the record the variance at each frequency occurred. A signal whose frequency changes over time, or whose amplitude comes and goes, will appear as a smear of power across a range of frequencies, indistinguishable from a signal containing those frequencies throughout. For stationary processes this averaging is exactly what is wanted. For non-stationary ones it can obscure the most interesting behaviour, which is the motivation for the wavelet methods of Chapters 5 and 6. The periodogram also treats the data as if they came from a process whose spectrum is a meaningful, fixed object. For a stationary random process, it is: the spectral density describes the process, and the periodogram of any particular record is a noisy estimate of it. How noisy, and what to do about it, is the subject of the next chapter. Chapter 3. Estimating Spectra Without Fooling Yourself The periodogram has a property that surprises almost everyone who first meets it. As the record gets longer, it does not get more accurate. Double the length of a record of white noise and compute its periodogram again: the values are spaced twice as closely in frequency, but at each frequency they scatter just as widely around the true flat spectrum as before. The standard deviation of each periodogram value is roughly equal to its expected value, regardless of how much data is used. In the language of statisticians, the raw periodogram is an inconsistent estimator of the spectral density. The reason is simple once seen. The periodogram at each Fourier frequency is computed from just two numbers, the cosine and sine coefficients at that frequency. For a Gaussian process, each periodogram value is therefore distributed as the true spectrum multiplied by a chi-squared variable with two degrees of freedom, divided by two: an exponential distribution. Adding more data adds more frequencies but not more information per frequency. The periodogram of a long record looks like a dense thicket of spikes, and some of those spikes will be large purely by chance. This chapter is about turning the periodogram into a statistically useful estimate of the spectrum, and about the harder problem that follows: deciding whether a peak in that estimate reflects something real. Trading resolution for stability Every practical remedy for the periodogram's instability works the same way. It averages several roughly independent estimates at each frequency, trading frequency resolution for reduced variance. The methods differ in how they obtain the independent estimates. Smoothing across frequency, sometimes called the Daniell method after Percy Daniell's 1946 proposal, averages the periodogram over a band of adjacent Fourier frequencies. Averaging M adjacent values yields an estimate with approximately 2M degrees of freedom and a variance reduced by a factor of M, at the cost of a resolution M times coarser. Because adjacent periodogram values of a stationary process are nearly independent, this works well for smooth spectra. It works poorly near sharp peaks, which it spreads out, and near steep spectral slopes, where the leakage of the underlying periodogram contaminates the average. The lag-window or Blackman–Tukey method, dominant before the FFT, estimates the autocovariance function, down-weights its values at long lags with a lag window, and Fourier-transforms the result. Down-weighting long lags is mathematically equivalent to smoothing the periodogram with a kernel, so this method is closely related to the Daniell approach. It is now mainly of historical interest, but it established the vocabulary of bandwidth and degrees of freedom still in use. Segment averaging, the method introduced by Peter Welch in a 1967 paper in the IEEE Transactions on Audio and Electroacoustics, splits the record into shorter segments, windows each one, computes its periodogram, and averages across segments. Welch showed that overlapping segments by half their length, with a Hann-type window, recovers much of the information lost to tapering at segment ends. With K segments, the variance falls by roughly a factor of K, while resolution coarsens to the inverse of the segment length. Welch's method is simple, robust and fast, and it is the default spectral estimator in many software packages. Its main weakness is that frequencies below the inverse of the segment length are lost entirely, a serious sacrifice when the low-frequency end of a geophysical spectrum is the point of interest. The choice among these methods is less important than the recognition that every spectral estimate embodies a bias–variance trade-off, controlled by a bandwidth parameter, whether the number of smoothed frequencies, the lag-window truncation, or the segment length. A narrow bandwidth gives high resolution and high variance; a broad one gives low variance and blurs features. There is no universally correct setting. The bandwidth should always be reported, and a spectrum should never be interpreted at a finer frequency scale than its bandwidth supports. Thomson's multitaper method In 1982 David Thomson of Bell Laboratories published "Spectrum estimation and harmonic analysis" in the Proceedings of the IEEE, a long and dense paper that introduced what is now called the multitaper method. It has become one of the most widely recommended spectral estimators in geophysics, and Donald Percival and Andrew Walden's 1993 textbook Spectral Analysis for Physical Applications made it accessible to a broad audience. The multitaper idea addresses the leakage–variance trade-off directly. Instead of cutting the record into segments, it uses the full record several times, each time with a different taper. The tapers are the discrete prolate spheroidal sequences, or Slepian sequences, studied by David Slepian and colleagues at Bell Labs in the 1960s and 1970s. For a chosen bandwidth W, these sequences are the functions of finite length whose Fourier transforms are most concentrated within the frequency band from −W to W. The first sequence is the best possible taper for concentrating energy in that band; the second is the best possible taper orthogonal to the first; and so on. Roughly the first 2NW of them, where N is the number of samples and W is the half-bandwidth in cycles per sample, have excellent concentration. The product NW, the time–bandwidth product, is the single parameter the analyst chooses, with values of 2 to 4 being typical. Because the tapers are orthogonal, the spectra computed with each are approximately independent, and averaging K of them yields an estimate with about 2K degrees of freedom. Because each taper is optimally concentrated, leakage from outside the band is very small. And because each taper emphasises different parts of the record, with higher-order tapers giving more weight to the ends, the method recovers information that a single taper throws away. Thomson also introduced an adaptive weighting scheme that down-weights the higher-order tapers at frequencies where the spectrum is steep and their leakage would be damaging. Two further features make multitaper estimation especially useful in physical sciences. First, it provides a natural statistical test for line components, meaning truly periodic signals such as tides, orbital periods, or instrumental artefacts at fixed frequency. Thomson's harmonic F-test compares the power explained by a sinusoid at each frequency with the residual power in the surrounding band, yielding a test statistic whose distribution under the null of no line is known. This separates the question "is there a periodic component here?" from the question "what is the shape of the continuous background?" Second, the approximately independent eigenspectra allow uncertainty to be estimated through the jackknife, by recomputing the estimate with each taper left out in turn, which gives confidence intervals that do not rely on distributional assumptions. Multitaper methods have limitations. The bandwidth parameter still trades resolution against variance, and a poor choice can hide closely spaced features or leave excessive noise. The harmonic F-test assumes a line component of constant amplitude and phase throughout the record, which is inappropriate for quasi-periodic signals like the solar cycle or ENSO. And the method as originally formulated assumes regular sampling. Within those constraints, it is the method against which others are usually judged. Table 2 summarises how the main non-parametric estimators compare on the criteria that matter most for physical records. Table 2. Non-parametric spectral estimators compared. Estimator How variance is reduced Leakage control Low-frequency coverage Main weakness Raw periodogram Not reduced Poor unless tapered Full record Inconsistent; spurious peaks Smoothed periodogram (Daniell) Averaging adjacent frequencies Depends on taper Full record, but lowest bands biased Smears sharp features Welch segment averaging Averaging overlapping segments Good (tapered segments) Lost below 1/segment length Sacrifices low frequencies Multitaper (Thomson) Averaging orthogonal tapers Excellent Full record Bandwidth choice; assumes even sampling Parametric AR (Burg, Yule–Walker) Model constraint Implicit Full record Model order choice; spurious splitting The last row anticipates Chapter 4, where autoregressive models are shown to define spectra of their own. What significance means for a spectral peak Suppose the spectrum of a century of annual temperature reconstructions shows a peak near a period of 22 years. Is it real? The question sounds straightforward but contains several separate questions, and much of the confusion in the literature arises from conflating them. The first question is what the null hypothesis is. A peak is significant only relative to some model of what the spectrum would look like without it. For most geophysical records, as Chapter 1 argued, a white-noise null is inappropriate, because the continuous background is red. The standard alternative, formalised for climate data by Keith Gilman, Frank Fuglister and John Mitchell in 1963 and made routine in the geosciences in the following decades, is to fit an AR(1) model to the data, compute its theoretical spectrum, and ask whether the observed spectrum exceeds it by more than chance would allow. The AR(1) spectrum has a simple form determined by two parameters: the variance of the process and its lag-one autocorrelation, often written ρ or α. Given these, the expected spectrum at each frequency is known, and because a smoothed spectral estimate with ν degrees of freedom is distributed approximately as the true spectrum times a chi-squared variable with ν degrees of freedom divided by ν, a confidence level can be drawn above the AR(1) curve. Peaks rising above, say, the 95 per cent line are candidates for significance. The second question is how the null was estimated. The AR(1) parameters are usually estimated from the same data, and if the data contain a strong oscillation, that oscillation will inflate the estimated variance and alter the lag-one autocorrelation. Mann and Lees in 1996 proposed estimating the background by fitting the red-noise model to a median-smoothed version of the spectrum, reducing the influence of peaks on the fit. Their method was later criticised for being too permissive in some settings, which illustrates that there is no neutral way to estimate a background: any procedure makes assumptions about what the peaks look like. The third question is how many frequencies were examined. A 95 per cent confidence level means that, under the null, each frequency has a 5 per cent chance of exceeding it. A spectrum with a hundred independent frequency bands will, under the null, show about five exceedances somewhere. Reporting the one peak that crosses the line, without accounting for the others that might have, is the multiple testing problem, and it is responsible for a large share of spurious periodicities. The simplest remedy is a Bonferroni-type correction; more refined approaches estimate the distribution of the largest peak across all frequencies under the null. If a peak was predicted in advance, for example at a known astronomical frequency, then a single test at that frequency is legitimate. If it was found by searching, the search must be accounted for. The fourth question is whether the peak is robust. Does it survive a change of taper, bandwidth, or detrending method? Does it appear in both halves of the record? Does it appear in independent records of the same phenomenon? A peak that depends delicately on analysis choices is not evidence of much, however high it rises above a confidence line. Monte Carlo testing An increasingly common approach to significance avoids analytic distributions entirely. One fits the null model to the data, generates many synthetic series from it with the same length, sampling, gaps and preprocessing as the real record, subjects each to exactly the same analysis, and compares the real result with the distribution of synthetic results. This is Monte Carlo or surrogate-data testing. Its great virtue is honesty. Every choice that affects the result, from detrending to tapering to handling of gaps, is applied equally to the real and synthetic data, so the test automatically accounts for them. It also accommodates null models for which no analytic spectrum is available, such as power-law noise, or nulls that preserve features of the data other than the one being tested. In the variant known as phase randomisation, surrogates are constructed by taking the Fourier transform of the data, randomising the phases while keeping the amplitudes, and transforming back. This preserves the entire power spectrum and hence the autocorrelation, and is useful for testing nonlinear structure or cross-relationships, though it cannot test a spectral peak, because it preserves that peak by construction. Myles Allen and Leonard Smith applied the Monte Carlo idea to singular spectrum analysis in a 1996 paper in the Journal of Climate, showing that many oscillations previously claimed in climate records failed to beat a properly constructed red-noise ensemble. Their framing, that the null should be tested with the same machinery as the data, has become standard practice. Three cautionary cases The history of spectral analysis in physical sciences offers instructive examples of both success and failure. The first success belongs to paleoclimatology. In 1976 James Hays, John Imbrie and Nicholas Shackleton published "Variations in the Earth's orbit: pacemaker of the ice ages" in Science. Analysing oxygen-isotope and other records from two Southern Ocean sediment cores spanning roughly the last 450,000 years, they found spectral peaks near periods of 23,000, 42,000 and about 100,000 years, matching the periods predicted for precession, obliquity and eccentricity of the Earth's orbit by the astronomical theory associated with Milutin Milanković. The finding held up because it tested predicted frequencies rather than searching for any peak, because the peaks were found in independent records, and because they were subsequently reproduced in many other cores. It remains a model of how spectral evidence can support a physical theory. Yet even this success contains a warning. The age model of the cores depended partly on assumptions about sedimentation rate, and later work showed that tuning age models to orbital curves can manufacture the very periodicities being sought. The dominance of the 100,000-year cycle, whose astronomical forcing is weak, remains a subject of active research. Spectral peaks can be robust while their interpretation remains contested. The second case comes from the long search for periodicities in the occurrence of earthquakes. Claims that large earthquakes cluster at particular phases of the tidal cycle, the lunar month, or the solar cycle have appeared many times. The most careful studies do find a small, statistically significant tidal modulation of earthquake occurrence in some settings, notably for shallow thrust earthquakes at times of large tidal stress, as reported by Elizabeth Cochran, John Vidale and Sachiko Tanaka in Science in 2004. But many earlier claims failed because they searched many periods and phases, used declustered catalogues inconsistently, or ignored the red character of earthquake occurrence produced by aftershock sequences. The effect, where real, is weak, and detecting it required precise prior hypotheses about the relevant stress. The third case is instrumental. Spectra of data from space missions and ground instruments commonly show sharp lines at frequencies associated with the instrument itself: spacecraft rotation, reaction-wheel speeds, thermal cycles, the mains power frequency, and the orbital period. The Kepler mission's data contained artefacts associated with the spacecraft's thermal environment and with its roughly quarterly roll, and the Transiting Exoplanet Survey Satellite shows systematics tied to its 13.7-day orbit. A harmonic F-test will rightly declare these significant. Whether they belong to the star or the spacecraft is not a statistical question. The lesson is that significance tells you that something is there, not what it is, and the most dangerous peaks are those that are both real and misattributed. Reporting a spectrum responsibly Good practice for reporting a spectral analysis in the physical sciences can be stated briefly. Report the estimator and its parameters: taper type, time–bandwidth product or segment length, overlap, and resulting bandwidth and degrees of freedom. Report the preprocessing: mean removal, detrending, prewhitening, gap handling. Plot the spectrum on logarithmic axes, so that power-law backgrounds appear as straight lines and confidence intervals have constant width, and show the bandwidth. State the null model, how it was estimated, and how multiple testing was handled. Show that key features are robust to reasonable changes in analysis choices. None of this is bureaucratic caution. Each item corresponds to a known way that a real-looking spectral feature can be an artefact, and a reader who cannot see the choices cannot judge the result. Hashtags: #TimeSeriesAnalysisForPhysicalSciences #TimeSeriesAnalysis #PhysicalSciences #SpectralAnalysis #SpectralDecomposition #FourierAnalysis #FastFourierTransform #PowerSpectrum #Periodogram #WaveletAnalysis #ContinuousWaveletTransform #WaveletCoherence #ARIMAModels #AutoregressiveModels #Stationarity #NonStationarity #Autocorrelation #RedNoise #SpectralLeakage #MultitaperMethod #LombScarglePeriodogram #IrregularSampling #SignalProcessing #MonteCarloTesting #FutureOfTimeSeriesAnalysis
- Fieldwork in Challenging Environments (Risk Management and Data Security)
Download the Book (PDF): Introduction On 25 January 2016, the fifth anniversary of the uprising that toppled Hosni Mubarak, a Cambridge doctoral student named Giulio Regeni left his flat in Cairo to meet a friend and did not arrive. His body was found by a road on the outskirts of the city nine days later, bearing the marks of prolonged torture. Regeni was studying independent trade unions among Cairo's street vendors, a subject that looked, from a British university, like ordinary labour sociology. In Egypt in 2016 it looked like something else. The investigations that followed, in Italy and in Britain, turned on questions that every researcher working in difficult places should be able to answer before departure and usually cannot: who knew what he was studying, who might have passed that information on, what his supervisors understood about the risk, and what anyone was supposed to do if he went quiet. Two years later, in May 2018, Matthew Hedges, a Durham doctoral student researching security policy in the United Arab Emirates, was detained at Dubai airport as he tried to leave the country. He spent months in solitary confinement and in November was sentenced to life imprisonment for espionage before being pardoned days later under intense diplomatic pressure. Part of the case against him rested on material taken from his own devices and notes. In September 2018 Kylie Moore-Gilbert, a Middle East specialist at the University of Melbourne, was arrested in Iran after attending an academic conference; she was held for more than two years before being released in a prisoner exchange. Xiyue Wang, a Princeton doctoral student working in Iranian archives on nineteenth-century history, spent more than three years in Tehran's Evin prison from 2016. None of them was doing anything that their home institutions would have described as dangerous. All of them were doing research in places where the meaning of research was decided by someone else. These cases are the ones that made the newspapers, and they are drawn from the social sciences. The quieter record is larger and more varied. Geologists work alone on unstable slopes, in artisanal mining districts where a stranger with a hammer and a GPS unit may be taken for a prospector or a government surveyor, and in borderlands where the rocks worth studying sit under disputed territory. Ecologists spend months in protected areas that are also contested land, where park guards, poachers, loggers, settlers and armed groups all have reasons to want to know who is counting what. Anthropologists live inside communities, which means their notebooks and recordings contain precisely the information that people in power would most like to have. The hazards differ by discipline, but the structure of the problem does not: a researcher enters a place where they are less informed, less connected and less protected than almost anyone around them, collects information that has value to others, and must eventually leave with it, while the people who helped them stay behind. What this book argues This book makes one argument and follows it through the practical work of fieldwork. The argument is that personal safety, data security and the protection of local collaborators are not three problems but one. Each is a question about how risk is distributed among the people involved in a project, and each is decided by the same choices: what the research asks, where and when it is done, who is involved, what is recorded, how information moves, and how the project ends. Treat them separately and you get the familiar pattern of modern fieldwork: a risk assessment form completed to satisfy an insurer, a vague intention to "back things up", and an unexamined assumption that the local research assistant will look after themselves. Treat them together and they become what they should always have been: part of the research design, decided early, argued over with the people who will carry the consequences, and revisited as conditions change. That reframing has practical consequences. It means that risk management starts with the research question, not with the travel booking. It means that a communication plan and a data plan are written by the same person on the same day, because the phone that carries your check-in messages also carries your interview recordings. It means that local collaborators are consulted about the threat picture as experts, paid and credited as colleagues, and planned for after the researcher leaves. And it means that the most important security decision in most projects is not which encryption software to use but what to collect in the first place. Who this book is for The book is written for researchers in the field sciences and the field-based social sciences who work, or are about to work, in remote or politically unstable places: doctoral students preparing a first long field season, supervisors and principal investigators responsible for other people's safety, and the research office staff who approve and insure such trips. It assumes intelligence and seriousness but no specialist background in security management or information technology. Geologists, ecologists and anthropologists appear throughout because their work covers the full range of what "challenging environments" means, from the physically remote to the politically hostile, and because each discipline has habits that the others would benefit from borrowing. Geologists tend to be good at physical hazard planning and weak on the politics of access. Anthropologists think hard about relationships and confidentiality and often neglect the mechanics of emergency communication. Ecologists are frequently expert at long remote deployments and underestimate how sensitive their location data can be. The book does not replace institutional training, a hostile environment course, a wilderness first aid qualification, or specialist advice for kidnap, detention or evacuation. It explains how those pieces fit together and what questions to ask of them. Where it describes tools, it describes categories and principles with named examples, because specific products change faster than books do and because the principles are what let a researcher judge a tool they have never seen before. How the book is organised The eight chapters follow the arc of a project. Chapter 1 sets out the case for treating risk as a design variable, drawing on the legal idea of duty of care and on what went wrong when institutions treated it as paperwork. Chapter 2 turns to reading a field site: building a picture of the actors, hazards and warning signs that matter, and deciding in advance what would make you stop. Chapter 3 is about local collaboration, the single largest determinant of both safety and data quality in difficult places, and about the risks that research assistants, drivers, translators and hosts carry on a project's behalf. Chapters 4 and 5 deal with communication. The first covers the routine: the check-in schedule, the equipment that carries it, and the overdue procedure that turns a missed call into action. The second covers the crisis: what happens in the first hours after an incident, who speaks to whom, how families, employers, embassies and the media are handled, and how the decision to evacuate or stay put is made. Chapters 6 and 7 turn to data. Chapter 6 builds a threat model for field data: who might want it, how they could get it, and what harm would follow for whom. Chapter 7 translates that model into practice, with encryption, backup routines and device discipline that work in places with unreliable power, poor connectivity and hostile border officials. Chapter 8 deals with endings: leaving the field, protecting collaborators and participants after departure, deciding what to publish and what to withhold, archiving or destroying data, and attending to the researcher's own recovery. The conclusion draws out what the argument implies for institutions as well as for individual researchers. A note on evidence The literature on field safety is uneven. There is a substantial body of reflective writing by anthropologists and political scientists who have worked in violent settings, a growing set of practical guides from the humanitarian and journalism sectors, a small number of systematic surveys, and a great deal of institutional guidance of variable quality. Hard statistics on incidents affecting researchers are scarce, because universities rarely collect or publish them and because the most serious incidents are often handled confidentially. Where this book cites a figure, it comes from a named source. Where evidence is thin, it says so and argues from principle and from documented cases instead. The aim throughout is to give researchers the judgement to make good decisions in situations no guide anticipated, which is, in the end, the situation that fieldwork in challenging environments always produces. Chapter 1. Risk as Method Most universities now require a risk assessment before a researcher travels to a remote or unstable place. The typical form asks the applicant to list hazards, rate each for likelihood and severity, describe control measures, and sign. It is reviewed by a department safety officer or a travel office, sometimes by an insurer, and filed. For a great many projects this is the whole of the institution's engagement with risk, and it happens at the end of the planning process, after the research question, the site, the methods, the budget and the timetable have been fixed. This sequence has the logic backwards. By the time the form is completed, nearly every decision that determines how dangerous the project will be has already been made. The question decides whom the researcher must talk to and what they must ask. The site decides the political and physical environment. The methods decide how long the researcher stays, how visible they are, and what they record. The budget decides whether there is money for a second vehicle, a satellite communicator, a trusted driver, or a local partner paid well enough to say no when something is wrong. The timetable decides whether fieldwork overlaps an election, a rainy season or a harvest. A risk assessment written after all of that can only list hazards the design has already created and suggest precautions around the edges. The argument of this chapter is that risk belongs inside research design, as a variable that is weighed alongside feasibility, originality and cost from the first draft of a proposal. That is not a counsel of caution. Treating risk as a design variable often makes ambitious research more possible, because it lets researchers see which elements of a project carry most of the danger and change those elements rather than abandoning the whole. Duty of care and what it actually requires The organising legal and ethical idea is duty of care: the obligation of an organisation to take reasonable steps to protect people who work for it or on its behalf from foreseeable harm. The concept is old in employment law, but its application to people working in dangerous places abroad was sharpened by a case from the humanitarian sector rather than from academia. In June 2012, gunmen attacked a convoy carrying staff of the Norwegian Refugee Council in the Dadaab refugee camp complex in north-eastern Kenya. A Kenyan driver was killed, and four international staff, including a Canadian named Steve Dennis, were abducted and taken across the border into Somalia, where they were freed in an armed rescue operation four days later. Dennis, who was wounded during the attack, later sued the organisation in Norway. In November 2015 the Oslo District Court found that the Norwegian Refugee Council had been grossly negligent and awarded him compensation of around 4.4 million Norwegian kroner. The court's reasoning, examined at length in a review by the European Interagency Security Forum (Merkelbach and Kemp, 2016), turned on whether the risk had been foreseeable and whether the organisation had taken measures proportionate to it. The court found that the organisation had information pointing to a heightened threat of abduction of foreigners in the area, that it had not put in place the protective measures that this threat called for, and that it had not adequately informed staff of the risks they were accepting. The case did not create new law, but it clarified what an organisation's duty of care means in practice, and universities are organisations with the same exposure. Three elements are worth drawing out. First, the standard is foreseeability, not certainty: an institution is expected to act on credible information about risk even when the precise incident cannot be predicted. Second, the measures must be proportionate and actually implemented; a plan that exists on paper but is not followed offers no protection and, arguably, makes the institution's position worse, because it shows that the risk was understood. Third, informed consent applies to employees and researchers as well as to research participants: people cannot meaningfully accept a risk that has not been explained to them. For academic fieldwork the duty of care is complicated by the ambiguous status of many researchers. Doctoral students are often not employees. Postdoctoral researchers may be employed by one institution and hosted by another. Local research assistants are frequently hired on short, informal arrangements, paid in cash, and invisible to the institution's insurance and safety systems altogether. The ethical duty does not disappear because the contractual relationship is unclear. If a project cannot be done without people who fall outside the institution's formal protection, the design needs to account for them explicitly. Chapter 3 returns to this at length. The problem with the standard risk form The likelihood-and-severity matrix at the centre of most institutional forms is not useless, but it encourages three distortions that matter in challenging environments. The first is that it treats hazards as independent items. A form lists "road traffic accident", "malaria", "robbery" and "political unrest" on separate rows, each with its own mitigation. In practice the serious incidents in fieldwork are usually chains. A researcher delays departure from a village because an interview runs long, drives on an unfamiliar road after dark, meets an improvised checkpoint, cannot reach anyone because the phone has no signal, and is carrying a notebook full of names. No single link is rated high on the form. The chain is what kills people or lands them in detention, and it is invisible to a method that scores items one at a time. The second distortion is that the matrix measures risk to the researcher and ignores risk created by the researcher for others. A form rarely asks what happens to an interviewee if the recording is seized, or to a driver if the vehicle is stopped and the passenger's documents are found to be incomplete, or to a village if a geological survey is taken as a sign that a mining concession is coming. These risks are real, frequently more severe than the risks to the researcher, and fall on people with less protection and fewer options. The third distortion is that the form is static. It is completed once, before departure, on the basis of information that may be weeks or months old. Fieldwork in unstable places is defined by change. A risk assessment that is not revisited when an election is called, a road is closed or a colleague is questioned by the police is a record of what the researcher once thought, not a tool for managing what is happening. The international standard for risk management, ISO 31000, first published in 2009 and revised in 2018, frames risk as "the effect of uncertainty on objectives" and describes risk management as an iterative process of establishing context, identifying, analysing and evaluating risk, treating it, and then monitoring and reviewing. The language is managerial, but the substance is useful for fieldwork precisely because it insists on context first and on repetition. The instrument that follows from it is not a form but a living document, which in this book is called the field risk plan: a short record of the threats that matter, the decisions that have been made about them, the indicators that would change those decisions, and the people responsible for acting. Who carries the risk A useful way to begin any field risk plan is to ask not "what could go wrong?" but "who carries the risk of each decision?" The question changes the analysis in ways that a hazard list does not. Take a decision as mundane as where to store interview recordings. If they are kept on the researcher's laptop without encryption, the researcher carries some risk, of losing their data, of embarrassment, perhaps of expulsion from the country if the laptop is inspected. The interviewees carry a different and often much larger risk: identification, questioning, loss of employment, violence. The research assistant who introduced the researcher to those interviewees carries a third kind of risk, because they are the known link between the foreigner and the community and they will still be there next year. The home institution carries reputational and legal risk. The same decision distributes different harms to different people, and the people who carry the heaviest harms are usually the ones with least say in the decision. The same logic applies to physical hazards. A geologist who decides to work alone on a remote outcrop to save the cost of a field assistant carries most of that risk personally. But if the geologist is injured and a rescue is required, the risk is transferred to the local search team, the helicopter crew, and the villagers who carry a stretcher down a hillside at night. An ecologist who stays in a protected area after an armed group has been reported nearby exposes not only themselves but the rangers who will be expected to find them if they go missing. Risk is never only personal, even when it feels like a private choice. Asking who carries the risk does three things. It makes the invisible participants in fieldwork visible to the planning process. It reveals where risk can be reduced without sacrificing research quality, because many decisions that load risk onto others are made for convenience rather than necessity. And it sets up a conversation with local collaborators in which their view of the risk is solicited as expert knowledge rather than as a courtesy, since they are often the only people in the planning process who know what the consequences look like on the ground. How the disciplines differ Although the logic is shared, the texture of risk differs by discipline, and researchers tend to inherit their field's blind spots along with its strengths. The contrasts are worth setting out because they show what each tradition can learn from the others. Table 1 compares the three disciplines that appear most often in this book across the dimensions that most shape their risk profiles. Table 1. Typical risk profiles of three field disciplines. Dimension Field geology Field ecology Field anthropology Typical setting Remote terrain, mines, borderlands Protected areas, forests, wetlands Villages, towns, informal settlements Dominant physical hazards Falls, rockfall, heat, cold, isolation Wildlife, disease, drowning, isolation Traffic, disease, crime Dominant human hazards Suspicion over resources, land disputes Poaching networks, armed groups, land conflict State surveillance, informant exposure Most sensitive data Precise locations, resource indications Species locations, patrol and poaching data Identities, testimonies, networks Typical visibility Low to moderate, brief visits Low, long residence in remote camps High, long residence among people The table simplifies, as any such comparison must. An anthropologist studying pastoralists may face every physical hazard a geologist does, and an ecologist working on fisheries may spend as much time in villages as any ethnographer. But the pattern is recognisable. Geologists are trained from their first field course to think about physical hazard: they learn to carry a first aid kit, leave a route plan, check the weather and avoid working alone on steep ground. Their characteristic blind spot is the human meaning of their presence. In districts where land and mineral rights are contested, a team taking samples and recording coordinates can be read as the advance party of a mining company, a government boundary survey, or foreign interests. The geological literature on field safety has historically had much less to say about this than about rockfall. Ecologists often run the longest and most remote deployments, and their institutions have built considerable expertise in logistics, medical preparation and camp management. Their blind spot tends to be the sensitivity of their data. Location records of rare, valuable or poached species are among the most commercially desirable information in the natural sciences; David Lindenmayer and Ben Scheele argued in a short and influential piece in Science in 2017 that ecologists should often decline to publish precise locations because collectors and poachers were mining scientific papers and databases for them. Data on ranger patrols and poaching incidents, which ecologists increasingly collect in collaboration with park authorities, can be dangerous in the hands of the networks they describe. Anthropologists have the richest tradition of reflection on danger, confidentiality and relationships, from the collection Fieldwork Under Fire edited by Carolyn Nordstrom and Antonius Robben (1995) to Lee Ann Fujii's work on the relational ethics of interviewing in violent settings. Their characteristic blind spot has been the practical machinery: communication schedules, evacuation planning, encryption. An ethnographer may think with great subtlety about what it means to record a testimony and still carry that recording unencrypted through an airport where devices are routinely examined. Building risk into the proposal If risk is a design variable, it should appear in the proposal, and the most effective place for it is not a separate section but inside the methods. Several practical habits follow. Begin by identifying the decisions in the design that carry most of the risk, and asking of each whether it is essential to the research question. Researchers are often surprised how many are not. A political ecologist who planned to interview both park authorities and residents accused of illegal grazing might find that the residents' perspective can be reached through group discussions organised by a local association, which carries far less risk for individuals than one-to-one interviews initiated by a foreigner. A geologist who planned to sample across a disputed boundary might find that the scientific question can be answered from outcrops on one side, with remote sensing to fill the gap. These are not compromises of rigour; they are the ordinary trade-offs that every design involves, made explicit. Then budget for risk reduction as a research cost, not an overhead. The items that make the greatest difference to safety in difficult places are expensive and unglamorous: a reliable vehicle with a trusted driver, a second person in the field, a satellite communicator with a service plan, travel insurance with medical evacuation cover, a hostile environment course, adequate pay for local collaborators, time for relationship building before data collection begins, and contingency days that allow a researcher to wait out trouble rather than drive through it. A proposal that does not fund these is a proposal that expects someone to take risks for free, and that someone is rarely the principal investigator. Write decision points into the timetable. Many projects in unstable places will face a moment when conditions change: an election is called, a curfew imposed, a road cut, a colleague arrested. A design that anticipates such moments, with alternative sites, alternative methods or a planned pause, lets the researcher respond without the pressure of watching their funding and their thesis evaporate. Researchers who have no fallback tend to stay too long, because leaving means failure; researchers who have planned for interruption can leave early and still finish. Finally, involve the people who know. That means colleagues who have worked in the area recently, local academic partners, and wherever possible the people who will work alongside the researcher in the field. It also means institutional specialists, who in larger universities now include travel security advisers and information security staff. The goal is not to have the proposal approved but to have it challenged by people who can see what the researcher cannot. Ethics review and the gap it leaves Research ethics committees, institutional review boards and their equivalents have become the main institutional checkpoint for research involving people, and they are the natural place to integrate risk to participants and to the researcher. In practice they often do neither well. Many committees were designed around biomedical research, where risks are clinical and data security means compliance with health privacy law. Their templates ask whether data will be stored on a password-protected computer, which is almost irrelevant to the threats a researcher faces at a hostile border, and rarely ask what happens to participants if the researcher is detained. The gap has two consequences. Researchers learn to write ethics applications as compliance documents, describing the protections the committee expects rather than the ones the context requires. And committees, lacking regional or security expertise, cannot judge whether a proposed protection is adequate, so they tend either to approve plans that are dangerously thin or to reject projects outright on the basis of a country's reputation. Neither serves research in difficult places. A better practice, adopted by some institutions, is to route projects in high-risk contexts through a joint review in which ethics, travel security and information security staff sit together with the researcher and, ideally, someone with recent experience of the region. The field risk plan becomes the shared document for that review, and it is revisited when conditions change. Researchers whose institutions do not offer such a process can create a version of it themselves. Share the field risk plan with a supervisor, a colleague with regional experience, and the institution's travel or security office if one exists. Ask each of them the same question: what is the most likely way this project harms someone, and what have I missed? The answers are rarely comfortable, which is the point. The shape of a field risk plan A field risk plan need not be long. The best are short enough to read in ten minutes and to update in the field. Its elements, which the following chapters develop, are these: a brief statement of the research and why it requires this site and these methods; a context analysis identifying the actors, hazards and trends that matter; a list of the specific threats judged most serious, with the people who would bear each; the measures adopted against each and who is responsible; the indicators that would trigger a change of plan, and what that change would be; the communication plan and the overdue procedure; the data plan, including what will not be collected; the arrangements for local collaborators, including pay, insurance and what happens if they are threatened; and the exit plan. The document should name the people in the home institution who will act if something goes wrong and should be accessible to them at all times. None of this is exotic. It is the ordinary discipline of thinking through a project before starting it, extended to the people and information the project touches. What makes it hard is not complexity but habit: the habit of treating safety as something that happens after the real planning is done. The remaining chapters take each element in turn, beginning with the question on which all the others depend, which is how to read a place accurately enough to know what the risks are. Chapter 2. Reading the Ground A researcher's safety in a difficult place depends less on equipment or procedures than on the accuracy of their picture of the place. Almost every serious incident in fieldwork involves a moment when someone misread the situation: took a new checkpoint for a routine one, assumed that a friendly official spoke for the whole of the state, mistook a lull in violence for its end, or failed to notice that their own presence had changed how people behaved. The procedures described in later chapters matter, but they depend on a prior act of judgement about what is going on. This chapter is about how to make that judgement well, before arrival and continuously afterwards. The humanitarian security literature calls this work context analysis, and it has developed a set of tools over three decades that researchers can borrow directly. The most influential text, Koenraad Van Brabant's Operational Security Management in Violent Environments, published by the Humanitarian Practice Network in 2000 and substantially revised in 2010, set out an approach built on understanding the actors and dynamics of a place before choosing a security strategy. The core insight transfers readily to research: security is not something a person carries with them, like a first aid kit, but a relationship between them and the environment, and the environment has to be understood before the relationship can be managed. Mapping the actors The first task is to identify who has power in the area where the research will take place, what they want, how they are likely to view the research and the researcher, and how they relate to one another. In most challenging environments the list is longer than it first appears. The state is rarely a single actor. Central ministries may issue a research permit while provincial authorities, the police, the intelligence services and the military each hold their own view of foreigners asking questions. In many countries the body that approves a research visa has no influence over the body that decides whether to detain a researcher, and a permit from one can be read as a provocation by another. Researchers who have worked in authoritarian settings frequently describe the experience of being welcomed by an academic partner and a ministry while being monitored, and sometimes interrogated, by security services who were never consulted. Beyond the state there are non-state armed groups, which may control territory, levy taxes at roadblocks, and run their own intelligence networks; criminal organisations, which in resource-rich areas often overlap with armed groups and with elements of the state; traditional and religious authorities, whose permission may matter more than any government document; companies, particularly in extractive industries, which may have their own security forces and strong views about who should be studying their concessions; non-governmental organisations, which may be allies, gatekeepers or competitors for local attention; and the community itself, which is never homogeneous and in which the researcher's hosts, interviewees and assistants all occupy particular positions. For each actor the useful questions are the same. What are their interests in the area? What is their capacity to help or harm? How are they likely to interpret the research, given its subject, its methods and the identity of the researcher? What do they already know about it, and how will they find out more? Who in the research team has a relationship with them, and what kind? The answers are rarely certain, but setting them out makes assumptions visible and shows where the gaps in knowledge are. A geologist working in an artisanal gold district, for example, might discover on reflection that the actors most interested in the survey are not the ministry that issued the permit but the local dealers who buy gold from diggers, the militia that taxes the mining sites, and the diggers themselves, all of whom may read a survey team as a threat to their livelihoods. Sources and their limits Information about a field site comes from many places, each with characteristic biases. Government travel advisories, such as those issued by the United Kingdom's Foreign, Commonwealth and Development Office and the United States Department of State, are the most widely consulted and the most widely misunderstood. They are designed for the general traveller, reflect diplomatic as well as security considerations, and are often coarse in geographic resolution. An advisory against all travel to a province may be triggered by conditions in one district; conversely, an absence of warnings does not mean an area is safe for someone asking sensitive questions. Advisories matter nonetheless, because insurers and universities often tie their cover and approval to them, and because travelling against official advice can void a policy at the moment it is most needed. Researchers should read them, understand their consequences for insurance and approval, and treat them as one input rather than as the analysis itself. Conflict event datasets have become a valuable resource. The Armed Conflict Location and Event Data project, known as ACLED, compiles reported political violence and protest events worldwide with dates and locations, and its data can show trends in a district over months or years that no single news report conveys. Its limitation is that it records reported events, which means that places with weak media coverage, or violence that is not reported, will be underrepresented. The International Crisis Group and similar organisations publish qualitative analysis of conflict dynamics that helps explain the patterns in the data. Humanitarian coordination bodies and, where they operate, NGO safety platforms such as the International NGO Safety Organisation produce security briefings that are often the most current and granular available, though access is usually limited to registered humanitarian organisations. The most valuable sources are people: researchers who have worked in the area recently, local academic colleagues, journalists, NGO staff, and above all the people who will work with the researcher in the field. Recent experience matters because conditions change; a colleague's account from five years ago may be dangerously out of date. Local knowledge matters because outsiders consistently misjudge both what is dangerous and what is not. But people are also partial. An academic partner may underplay risks to preserve a collaboration; an NGO worker may overplay them because their organisation's security rules are stricter than a researcher's need to be; a local assistant may be reluctant to contradict a foreign employer. Good context analysis triangulates, asks the same question in different ways of different people, and pays attention to what people are reluctant to say as well as what they say. Physical and environmental hazards The political reading of a place gets most of the attention in discussions of dangerous fieldwork, but for many researchers the physical environment is the more likely source of harm. Nancy Howell's survey of anthropologists for the American Anthropological Association, published in 1990 as Surviving Fieldwork, remains one of the few systematic attempts to document what actually happens to researchers in the field, and one of its enduring findings was how much of the harm came from illness, vehicle accidents and environmental hazards rather than from violence. The pattern is familiar to anyone who has worked in remote places. Roads kill more field researchers than armed groups do. Physical hazards also interact with political ones in ways that matter for planning. A rainy season that cuts roads also cuts evacuation routes. A malaria episode that would be routine in a town becomes an emergency in a camp three days from a clinic, and more so if the only road passes through a militia checkpoint that is closed at night. Heat, altitude and cold impair judgement at precisely the moments when judgement matters most. The practical implication is that the context analysis should include a calendar: seasons, holidays, elections, harvests, migrations, and any other predictable events that change either the physical environment or the political temperature. Fieldwork timed to avoid the worst overlaps is safer than fieldwork that relies on precautions to manage them. For geologists and ecologists in particular, the physical hazard assessment should include the specific features of the terrain and the work: slope stability and rockfall, river crossings, tides and flash floods, wildlife, the distance to the nearest point from which a vehicle or helicopter can extract a casualty, and the time that extraction would take in the best and worst conditions. The last figure, sometimes called the evacuation time, is one of the most useful single numbers in a risk plan, because it determines what level of medical capability must be carried in the field. A team that is four hours from a hospital needs basic first aid; a team that is two days from one needs someone trained in wilderness medicine and the equipment to stabilise a serious casualty for that long. Identity and exposure Risk in the field is not distributed evenly among researchers. Gender, sexuality, ethnicity, nationality, religion, age, disability and perceived wealth all shape how a researcher is seen and what dangers they face, and the same person may be protected by one aspect of their identity and exposed by another. The most systematic evidence concerns sexual harassment and assault, which the field sciences long treated as a private matter. In 2014 Kathryn Clancy, Robin Nelson, Julienne Rutherford and Katie Hinde published the Survey of Academic Field Experiences in PLOS ONE. Of the 666 field scientists who responded, 64 per cent reported having personally experienced sexual harassment in the field, and over 20 per cent reported having experienced sexual assault. Women trainees were the most frequent targets, and the perpetrators they reported were predominantly their superiors within the research team, not strangers at the field site. Fewer than 40 per cent of respondents recalled ever encountering a code of conduct at a field site, and fewer than a quarter recalled a sexual harassment policy. The survey was not a random sample of all field scientists, and its authors were careful about that limitation, but it transformed the conversation by showing that one of the most significant risks in fieldwork came from inside the research team. That finding has a direct implication for risk planning: a field risk plan that considers only external threats is incomplete. Team composition, supervision arrangements, sleeping and transport arrangements, codes of conduct, and above all a reporting route that does not run through the person who might be the problem, are safety measures as real as any satellite phone. They belong in the same document. Nationality and ethnicity shape risk in less obvious ways. A researcher with a passport from a country whose government is in dispute with the host state may be a more attractive target for detention as a bargaining chip, as several of the cases in the introduction illustrate. A researcher who shares an ethnic or religious identity with one side of a local conflict may be read as partisan regardless of their own views. A researcher from the region, returning to study their own society, may have better access and language but less protection: they may not be seen as a foreigner whose detention would cause diplomatic trouble, and their family may be reachable in ways that a foreign researcher's family is not. Dual nationals face particular difficulties, since some states do not recognise a second nationality and will deny consular access on that basis. None of this is a reason for any researcher to avoid any place. It is a reason for each researcher's risk assessment to be personal rather than generic. Acceptance, protection and deterrence Once the context is understood, the next question is what strategy to adopt. The humanitarian literature distinguishes three broad approaches, usually drawn as the points of a triangle. Acceptance means reducing threats by building relationships and legitimacy, so that the people who could cause harm choose not to. Protection means reducing vulnerability through measures such as secure accommodation, low profile, careful movement and good communications. Deterrence means posing a counter-threat, through armed escorts, legal action or political pressure. Researchers rely overwhelmingly on acceptance and protection. Deterrence is rarely available and usually counterproductive: an armed escort marks a researcher as a target, associates them with whichever force provides it, and destroys the trust on which most field research depends. Acceptance is the strategy most aligned with good research, because it depends on the same things good research does: clear explanation of the project, respect for local authority structures, genuine engagement with people's concerns, and a reputation for keeping promises. Its limitation is that it works only with actors who are willing to be engaged and whose attitudes can be shaped. It offers little protection against a predatory criminal group, an intelligence service that has already decided a researcher is a spy, or a random act of violence. Protection fills that gap. It includes the choices about profile that researchers make every day: whether to drive a branded vehicle or an ordinary one, whether to carry a laptop to interviews or leave it locked away, whether to publicise the research or keep it quiet, whether to stay in a hotel frequented by foreigners or with a local family. Low profile is often the right default for sensitive research, but it has costs. A researcher who is too discreet may be suspected of having something to hide, and may lose the protection that visibility and recognised affiliations can offer. The right balance depends on the actors. Against an intelligence service, a transparent, well-documented research presence with a recognised local partner may be the best protection. Against criminal kidnappers, invisibility may be. Indicators, triggers and deciding to stop The most practically valuable output of a context analysis is not a description of the current situation but a set of signs that the situation is changing, linked in advance to decisions. Security professionals call these indicators and triggers. An indicator is something observable that would suggest risk is rising: a new checkpoint, a change in the tone of official contacts, rumours about the researcher, the departure of other foreigners, an increase in local violence, a colleague being questioned. A trigger is a threshold at which a predetermined action follows: suspending interviews, moving to a safer location, contacting the home institution, leaving the area or the country. The reason to set triggers in advance is that people in the field are bad at judging gradual change. The literature on dangerous fieldwork is full of accounts of researchers who stayed too long because each day seemed only slightly worse than the one before. Psychologists have a name for part of this, normalisation, the process by which repeated exposure to a threat makes it seem ordinary. There are other pressures too: a researcher who has invested years in a project, whose funding runs out in months, and whose local collaborators depend on the work continuing, has every incentive to interpret ambiguous signs optimistically. Deciding in advance, when calm and away from those pressures, what would make them stop is a way of binding their future self to a judgement made with a clear head. Good triggers are specific, observable and few. "If the security situation deteriorates" is not a trigger. "If any member of the team is questioned by security services about the project" is. So is "if the road between the field site and the district capital is closed for more than 48 hours", or "if the government declares a state of emergency in the province", or "if we receive a direct threat". Each trigger should be linked to an action and to a person responsible for taking or authorising it. Some triggers should be decided by the researcher in the field; some should rest with a named person at home who can see the situation from a distance and who is not subject to the same pressures to stay. Triggers also need to be revisited. The context analysis is not a document written once before departure. It should be reviewed at planned intervals, weekly in volatile places, and whenever an indicator fires. A short written update, shared with the home contact, forces the researcher to articulate what has changed and makes it easier for someone at a distance to notice a drift that the researcher has normalised. Reading oneself The last element of reading the ground is the hardest: reading the effect of the researcher's own presence. Researchers change the places they study. They bring money, attention, questions and connections to the outside world. In a community divided by a land dispute, whom the researcher stays with, who is hired as an assistant, and whom the researcher interviews first all send signals that are read closely by people with a great deal at stake. A survey that asks about violence can reopen wounds, stir suspicion or put respondents at risk from those who do not want the violence discussed. A geological survey can raise hopes of development or fears of dispossession. An ecological study can be seen as the first step towards a new protected area and the evictions that often accompany one. Jeffrey Sluka, an anthropologist who worked among republican communities in Northern Ireland during the conflict, argued in Fieldwork Under Fire that managing danger in the field depended above all on understanding how one was perceived and actively managing that perception: explaining the research openly, avoiding association with parties to the conflict, and being alert to rumours. His advice was specific to his setting, but the principle generalises. A researcher's reputation in the field is part of the environment they must read, and it is partly within their control. Reading oneself also means being honest about one's own condition. Fatigue, illness, loneliness, frustration and fear distort judgement, and all are common in long field seasons. Researchers who are exhausted or isolated take shortcuts, stop checking in, and misread signs they would have caught when fresh. A context analysis that ignores the state of the person doing the analysis leaves out one of the most important variables. The communication routines described in Chapter 4 are partly a safeguard against this, since a regular conversation with someone outside the field site is one of the few reliable ways to notice that one's judgement has drifted. Before that, however, comes the relationship that shapes both safety and research in the field more than any other, which is the relationship with the local people who make the work possible. Chapter 3. Partnership as Protection Every field project in a difficult place depends on people who are not named on the grant. Research assistants find interviewees, translate, transcribe and explain. Drivers choose routes, read checkpoints and decide when a road is too dangerous. Hosts provide food, shelter and a social identity that makes a stranger comprehensible. Local academic partners supply permits, institutional cover and knowledge. Guides, porters, boat operators, village chiefs and park rangers make remote sites reachable. In ecological and geological work, local field assistants frequently know more about the terrain, the species or the rock than the visiting scientist does. These people are the single largest determinant of both the safety and the quality of fieldwork in challenging environments. They are also, very often, the people who carry the most risk and receive the least protection, pay and credit. This chapter argues that the two facts are connected: that the relationships which make research safe are the same relationships which make it good, and that treating local collaborators as colleagues rather than as hired help is not only an ethical obligation but the most effective security measure available to most researchers. The invisible workforce The dependence of field research on local intermediaries is as old as fieldwork itself, and so is the habit of leaving them out of the story. Townsend Middleton and Jason Cons, in an article in Ethnography in 2014, examined how research assistants had been written out of anthropology's accounts of itself, reduced to acknowledgements or omitted entirely, even though the knowledge the discipline produced depended on their labour and judgement. Similar observations apply to the natural sciences, where local field assistants who located specimens, identified species or guided survey teams have often gone unrecorded in the publications that followed. The voices of those intermediaries became more audible in 2019, when a group of researchers based in eastern Democratic Republic of the Congo, many of whom had spent years working as assistants and collaborators for foreign scholars, published a series of blog posts that became known as the Bukavu Series. Their accounts described being paid little and late, being sent alone into dangerous areas to collect data that foreign researchers considered too risky to gather themselves, being denied authorship on the work they had made possible, and being left without support when the risks they had taken materialised. The series was notable not because the experiences were new but because they were written by the people who had lived them, and because they described in detail a pattern that many foreign researchers had benefited from without examining. The pattern has a security dimension that is easy to miss. When a foreign researcher delegates the most dangerous parts of data collection to local staff, the project's risk profile does not fall; it is merely transferred to people who are less visible to the institution, less protected by its insurance, and less likely to be evacuated if something goes wrong. A field risk plan that shows low risk for the researcher and does not mention the assistant who is conducting interviews in a contested district is not a record of a safe project. It is a record of a displaced risk. Why local collaborators see what outsiders miss The strongest practical argument for genuine partnership is epistemic. Local collaborators know things that visitors cannot learn in a few weeks or months. They can distinguish a routine checkpoint from an unusual one, a neighbourly question from an informant's probing, a genuine threat from local bravado. They know which families are feuding, which officials are reliable, which roads flood, which rumours are circulating. They can often tell, from small changes in behaviour, that something is shifting long before any public event confirms it. That knowledge is only available to the researcher if it is shared, and it is only shared if the relationship makes sharing possible. An assistant who is paid by the day, who fears losing the job, who has been treated as a translator rather than a colleague, and who has learned that the foreign researcher does not like to be contradicted will often not say that an interview is a bad idea or that a route is unsafe. The power asymmetry silences the person best placed to give warning. Researchers who have worked in violent settings consistently describe the moment when an assistant or host told them, often indirectly, that they should leave, and many describe regret at having been slow to listen. Lee Ann Fujii, whose research on the Rwandan genocide involved extensive work with local assistants, wrote about the ethics of these relationships with unusual care, including in a 2012 article in PS: Political Science and Politics titled "Research Ethics 101: Dilemmas and Responsibilities". Among her points was that researchers owe their assistants the same ethical consideration they owe their participants, that assistants' safety must be weighed as seriously as the researcher's, and that the researcher's decisions about whom to interview and how can expose assistants to consequences the researcher never sees. The conclusion that follows is practical: build a relationship in which the assistant's judgement carries real weight in decisions, and in which saying "no" or "not today" is safe. What partnership means in practice Partnership is easily invoked and less easily practised. Several concrete commitments distinguish it from rhetoric. The first is fair pay, set in consultation with local colleagues rather than by the lowest acceptable rate. Fair pay reflects the skill involved, the risks taken, the opportunity cost of the work, and the fact that fieldwork is often seasonal and irregular. It is paid on time, in full, and by a method that does not itself create risk: carrying large amounts of cash to pay a team is a well-known hazard, and payments that pass through visible channels may draw attention from officials or armed groups. Pay should include compensation for days when fieldwork is suspended for security reasons, since an arrangement that pays only for days worked creates a direct incentive for the assistant to downplay risk. The second is insurance and medical cover. Most university travel insurance covers only the institution's own staff and students. Local collaborators are typically excluded, which means that if the researcher and the assistant are injured in the same vehicle accident, one may be evacuated by air to a modern hospital while the other receives whatever care is available locally at their own expense. Some institutions can extend cover or purchase separate policies; where they cannot, the budget should include a contingency for collaborators' medical costs, and the collaborator should know in advance what support they can expect. The discussion of this topic is often uncomfortable, which is exactly why it needs to happen before departure rather than in an emergency. The third is involvement in risk decisions. Collaborators should see and contribute to the field risk plan, including the triggers for suspending or ending work. They should have the explicit authority to stop an activity they judge unsafe, without having to justify it and without financial penalty. They should be part of the communication plan, both as people who check in and as people whose absence triggers a response. The fourth is recognition. When local collaborators contribute intellectually to research, by shaping questions, interpreting findings or producing data through their own expertise, that contribution should be reflected in authorship according to the same criteria that apply to anyone else. Where authorship is not appropriate, acknowledgement should be specific. There is a complication: in some settings, public association with a foreign research project, especially one on a sensitive subject, is itself a risk. The right answer to that complication is to ask the collaborator what they want, not to decide on their behalf that anonymity is safer. The fifth is continuity. A relationship that ends when the grant ends leaves collaborators exposed in several ways. They may be associated in local memory with a project that later becomes controversial. They may hold data or know about data that others want. They may have lost other work while employed. A commitment to stay in contact, to warn collaborators if publication might draw attention to them, and to help if they are threatened because of the project is part of a responsible design. Chapter 8 returns to this. Local academic partners Formal partnerships with universities and research institutes in the host country have become standard, and in many countries they are legally required for research permits. They bring clear benefits: institutional legitimacy, local expertise, access to permits and networks, and, in the best cases, a genuine intellectual collaboration that improves the research. They also bring their own risks, which are less often discussed. A local academic partner may be subject to pressures the visiting researcher cannot see. In authoritarian settings, universities may be closely monitored by security services, and academic colleagues may be required or pressured to report on foreign visitors. This need not reflect bad faith on the colleague's part; it may simply be a condition of their continued employment. A partner may also bear the consequences if the project causes offence: the visiting researcher can leave, but the partner institution may face investigations, loss of permits, or worse. The implication is not to distrust partners but to be clear-eyed about their constraints, to discuss openly what the project involves and what risks it might create for them, and to avoid putting them in positions where their obligations to the project and to their own state conflict. For natural scientists, partnership also has a legal dimension. The Nagoya Protocol on Access and Benefit-sharing, adopted under the Convention on Biological Diversity in 2010 and in force since 2014, requires that access to genetic resources be based on prior informed consent from the providing country and on mutually agreed terms for sharing benefits. Its implementation varies widely between countries, and many ecologists find the permit processes slow and opaque. But researchers who collect biological samples without the proper permits and agreements expose themselves to legal risk, expose their partners to accusations of complicity in biopiracy, and undermine the trust on which future research depends. Geologists face comparable issues with export permits for rock, mineral and fossil samples, which some countries treat as national heritage or strategic resources. Getting these arrangements right is part of security, not a bureaucratic distraction from it. Fixers, drivers and the problem of borrowed trust In many difficult places, researchers work with people who are not academic colleagues but brokers: fixers who arrange access and logistics, drivers who know the roads and the checkpoints, guides who know the terrain and the people who control it. The journalism literature has examined this role closely, notably in Colleen Murrell's study of fixers in international newsgathering, published in 2015, and much of what it describes applies to researchers. A fixer's value lies in their relationships: they can get the researcher through a checkpoint, into a meeting with a local commander, or onto a boat, because the people involved know and trust them. This is borrowed trust, and it has two consequences. The first is that the researcher is protected only as far as the fixer's relationships extend; outside that network, the protection vanishes. The second is that the fixer's reputation is on the line with every introduction. If the researcher behaves badly, asks the wrong question, or publishes something that offends, it is the fixer who will face the consequences with the people they introduced, long after the researcher has left. Choosing brokers therefore requires care. Recommendations from other researchers and journalists who have worked with someone recently are the best starting point. It is worth understanding a fixer's own position in local networks, including political and ethnic affiliations that might make them welcome in some places and dangerous in others. A driver who is trusted in one district may be a liability in the next. It is worth being explicit about what the researcher will and will not do, so that the broker does not make promises on the researcher's behalf that cannot be kept. And it is essential to listen when a broker says no: their reluctance to take a route or arrange a meeting is often the most accurate risk assessment available. When collaborators are threatened The hardest scenarios in field partnerships are those in which a collaborator is threatened, detained or harmed because of the project. They are more common than published accounts suggest, because they are often handled quietly to avoid making things worse. Planning for them begins with an honest conversation before fieldwork starts about what the researcher and the institution can and cannot do. A foreign researcher usually cannot protect a local collaborator from their own government. They may be able to provide legal support, pay for a lawyer, contact human rights organisations that monitor detentions, help with temporary relocation within the country, or, in extreme cases, support an application for protection abroad. Some of these options depend on funds, networks and institutional willingness that must be arranged in advance; none can be improvised quickly from abroad in the middle of a crisis. Organisations that support threatened scholars, such as Scholars at Risk and the Council for At-Risk Academics, can advise and sometimes assist, though their capacity is limited and their criteria specific. The communication plan should include collaborators, and the crisis procedures described in Chapter 5 should cover what happens if a collaborator, rather than the researcher, is the one in trouble. That includes who decides whether to raise a detention publicly, which in some cases helps and in others harms; who contacts the collaborator's family; and what happens to the data the collaborator holds. The collaborator's own preferences, where they can be known, should guide those decisions. It is worth asking in advance, for example, whether a collaborator would want their situation publicised if they were detained, and whom they would want contacted. Data held by collaborators A final dimension of partnership connects directly to the data security chapters that follow. Local collaborators frequently hold research data: recordings on their phones, transcripts on their computers, contact details of interviewees, photographs, survey responses collected on tablets. In many projects the collaborator's devices are the weakest point in the data security chain, not because collaborators are careless but because they have not been given the tools, training or equipment the researcher has. A research assistant who uses a personal phone to record interviews, stores transcripts on a shared family computer, and sends files to the researcher over an unencrypted messaging app is exposed, and so are the interviewees. Providing dedicated, encrypted devices, training in their use, a secure channel for transferring data, and a clear rule about what should be deleted and when is a modest cost that closes one of the largest gaps in most projects. It also communicates something about the relationship: that the collaborator's safety is taken as seriously as the researcher's. The same principle applies to data access and ownership. Collaborators who have contributed to data collection may have legitimate interests in using the data in their own research, and in many cases they should. But access should be deliberate, secure and agreed, not an accident of files left on a laptop. A data management plan that describes who holds what, for how long, and under what protection is part of a responsible partnership, and it is a tool for protecting collaborators as much as participants. The through-line of this chapter is simple: the people on whom fieldwork depends are not a logistical resource but colleagues whose knowledge, safety and interests are integral to the research. A project designed around that principle is more ethical, but it is also safer, because the people best placed to warn of danger are empowered to do so, and it produces better data, because the relationships that make people willing to speak honestly are the same ones that make them willing to raise an alarm. The next two chapters take up the question of how those warnings travel, in routine and in crisis. Hashtags: #FieldworkInChallengingEnvironments #FieldworkRiskManagement #ResearcherSafety #DataSecurity #FieldResearch #RiskAssessment #DutyOfCare #ContextAnalysis #FieldRiskPlan #ThreatModeling #ResearchEthics #ResearchSecurity #LocalCollaborators #ResearchPartnerships #EmergencyCommunication #CrisisManagement #EvacuationPlanning #OperationalSecurity #DigitalSecurity #DataProtection #Encryption #SensitiveResearch #HostileEnvironments #ResearcherDutyOfCare #FutureOfFieldResearch
- Causal Machine Learning (Combining Prediction and Treatment Effect Estimation)
Download the Book (PDF): Introduction A hospital wants to know whether a new discharge programme reduces readmissions. A retailer wants to know whether a discount brings back lapsed customers or merely subsidises people who would have returned anyway. A ministry of labour wants to know which unemployed workers benefit from a training course and which would do as well without it. Each of these organisations already owns a great deal of data and, increasingly, a team that can fit a gradient-boosted model or a neural network to that data in an afternoon. The models will predict readmission, repurchase and re-employment with impressive accuracy. None of them, on their own, will answer the question that was asked. The gap between those two things, between predicting an outcome and estimating what an intervention does to it, is the subject of this book. It is an old gap. Statisticians and econometricians spent most of the twentieth century building a vocabulary for it: potential outcomes, confounding, identification, instruments, the propensity score. Machine learning grew up largely on the other side of the gap, with a different vocabulary and different habits: held-out test sets, loss functions, regularisation, the relentless pursuit of lower prediction error. For a long time the two traditions regarded each other with a mixture of interest and suspicion. Econometricians saw black boxes that could not produce a standard error. Machine learners saw economists fitting linear regressions with a dozen hand-picked controls to data sets that could support a thousand. Over roughly the past fifteen years, a body of work has emerged that takes both sides seriously. It is usually called causal machine learning, and its central achievement is surprisingly specific. It does not make machine learning causal. What it does is show precisely where flexible predictive models can be inserted into a causal analysis, what goes wrong when they are inserted carelessly, and how to repair the damage so that the final estimate comes with a valid confidence interval. Three families of methods carry most of the weight: double or debiased machine learning, developed chiefly by Victor Chernozhukov and colleagues in econometrics; targeted maximum likelihood estimation, developed by Mark van der Laan and colleagues in biostatistics; and generalized random forests, developed by Susan Athey, Julie Tibshirani and Stefan Wager at the meeting point of the two. They grew up in different departments, publish in different journals and use different notation. They are much closer to one another than their literatures suggest. The argument of this book The controlling claim of what follows is that machine learning earns its place in causal inference only as a servant of a well-defined estimand and a defensible identification strategy, and that the methods which make this partnership work all rest on one idea: building the estimator around a score that is insensitive to small errors in the predictive models it depends on. Get the estimand and the identifying assumptions right, and flexible learners let you drop the fragile functional-form assumptions that have always haunted applied work. Get them wrong, and no amount of predictive accuracy will rescue the answer; if anything, a sophisticated model will make a wrong answer look more authoritative. This claim cuts in two directions, and the book pursues both. Against the enthusiast who believes that a sufficiently powerful model will discover causal structure in observational data by itself, it insists that causal conclusions are purchased with assumptions, and that those assumptions come from knowledge of how the data were generated, not from the data alone. Against the sceptic who regards machine learning as a fashionable distraction, it shows that the classical workhorse of observational research, a linear regression with controls, silently imposes assumptions that are frequently false, and that the new methods relax those assumptions in a principled, well-understood way, with inference that is as rigorous as anything in the textbook canon. The idea that unifies the three families deserves a sentence of preview here, because it recurs in almost every chapter. Every estimator in this book needs to learn some auxiliary functions from the data: how the outcome depends on covariates, how the probability of treatment depends on covariates, or both. These are called nuisance functions because they are not what we care about; they are tools for getting at the treatment effect. Machine learning is extremely good at estimating nuisance functions and extremely bad at estimating them without bias, because regularisation, the very device that makes it predict well, deliberately shrinks estimates away from the truth. The trick, which goes back to Jerzy Neyman in the 1950s and to the semiparametric statistics of the 1980s and 1990s, is to construct the estimate of the treatment effect so that small errors in the nuisance functions have no first-order effect on it. Double machine learning calls this Neyman orthogonality. Targeted learning calls it solving the efficient influence curve equation. Generalized random forests build it in by centring the data before growing trees. The language differs; the mathematics is essentially the same. What the book covers and what it leaves out The first chapter sets out the causal framework the rest of the book relies on: potential outcomes, the estimands most analyses target, and the assumptions that allow those estimands to be recovered from data. It is deliberately insistent on the point that identification comes before estimation. The second chapter explains what goes wrong when one simply plugs a machine learning model into a causal formula, working through the regularisation bias and overfitting bias that make naive approaches unreliable. The third develops the remedy: orthogonal scores, influence functions and double robustness, explained without more mathematics than an attentive reader can follow. The next two chapters turn the remedy into procedures. Chapter 4 describes double/debiased machine learning in practice, including cross-fitting, the main model variants and the choices an analyst must make. Chapter 5 describes targeted maximum likelihood estimation, the super learner that usually accompanies it, and how the targeting step differs from, and resembles, the debiasing step of double machine learning. The last three chapters address heterogeneity, the question of whom a treatment helps and by how much. Chapter 6 introduces the conditional average treatment effect and the family of meta-learners and Bayesian tree methods that estimate it. Chapter 7 is devoted to generalized random forests and the causal forest, the most widely used tool for this problem. Chapter 8 addresses the uncomfortable fact that heterogeneous effect estimates are much harder to validate than average effects, and describes the tools, from best linear projections to rank-weighted average treatment effects and policy learning, that make them usable. It also returns to the assumptions of the first chapter, examining overlap and unmeasured confounding in the high-dimensional settings where machine learning is most tempting and those assumptions are most strained. Several important topics are left aside or touched only lightly. Causal discovery, the attempt to learn causal graphs from data, is a distinct enterprise with its own literature and receives no treatment beyond a warning. Difference-in-differences, regression discontinuity and synthetic control designs have all acquired machine learning extensions, but they would each need a chapter of background, and the book concentrates on the setting where the three core methods are most fully developed: a treatment whose assignment is explained by measured covariates, with instrumental variables as the main extension. Time-varying treatments and dynamic regimes, where targeted learning is particularly strong, are mentioned where relevant but not developed. Deep-learning architectures designed specifically for treatment effect estimation, such as representation-balancing networks, are acknowledged but not surveyed; they have not displaced the methods covered here in applied work, and they fit into the same framework of nuisance estimation and orthogonal scoring. How to read it The book assumes a reader who is comfortable with regression, conditional expectations and the idea of a sampling distribution, and who has at least a passing familiarity with methods such as random forests, the lasso and gradient boosting. It does not assume training in semiparametric theory. Where formulas help, they are written out in plain notation and explained in words; where they do not help, they are left out. Notation is kept consistent across chapters: Y is an outcome, D or A a treatment (economists prefer the first, biostatisticians the second, and the book follows whichever literature it is discussing), X or W the covariates, and Y(1) and Y(0) the potential outcomes under treatment and control. Readers from economics will find the potential outcomes framework and the partially linear model familiar and may move quickly through the first chapter; the material on targeted learning and super learning in Chapter 5 is where their tradition is thinnest. Readers from epidemiology and biostatistics will be at home with double robustness and TMLE, and may find the econometric treatment of instruments and the partially linear model in Chapter 4 the most novel. Data scientists trained in machine learning will find the first three chapters the most important, because they explain why a model that validates beautifully on a held-out set can still produce a badly biased treatment effect. The discipline this book asks of its reader can be put simply. Before choosing a learner, name the estimand. Before trusting an estimate, state the assumptions under which it means what you want it to mean, and check the ones that can be checked. Before acting on an estimated pattern of heterogeneity, test whether it is there. None of this is new. What is new is that we now have estimators flexible enough to take the data seriously and principled enough to report honestly how uncertain they are. Chapter 1. Estimands Before Estimators Machine learning is organised around a single question: given inputs, what output should we expect? A causal question has a different shape. It asks what the output would be if we reached into the world and changed one of the inputs, holding fixed the process that generated everything else. The two questions coincide only under special conditions, and the whole of causal inference can be read as a disciplined account of what those conditions are and what to do when they fail. A familiar example makes the difference concrete. Across many health systems, patients who receive intensive care are more likely to die in the following month than patients who do not. A model trained to predict mortality will rightly learn that admission to intensive care is a strong predictor of death. Nobody concludes that intensive care kills people. Sicker patients are sent there, and sickness drives both the treatment and the outcome. The predictive relationship is real and useful; it is simply not the answer to the question "what would happen to this patient's chances if we admitted them?" The predictive model has no way of telling the two apart, because the difference lies not in the data but in the process that produced them. Potential outcomes The most widely used language for making causal questions precise is the potential outcomes framework, associated with Jerzy Neyman's work on agricultural experiments in the 1920s and developed for observational studies by Donald Rubin from the 1970s onward. Its central move is to imagine, for each unit, a set of outcomes indexed by the treatment the unit might receive. For a binary treatment, each unit i has two potential outcomes: Yᵢ(1), the outcome it would experience if treated, and Yᵢ(0), the outcome it would experience if not. The individual causal effect is the difference Yᵢ(1) − Yᵢ(0). The difficulty, which Paul Holland in 1986 called the fundamental problem of causal inference, is that we never observe both. A patient is either admitted or not; a customer either receives the discount or does not. What we observe is Yᵢ = Dᵢ·Yᵢ(1) + (1 − Dᵢ)·Yᵢ(0), where Dᵢ records which treatment was actually received. Half of the information needed to compute any individual effect is always missing. Causal inference is, in this sense, a missing data problem, and many of the tools in this book were first developed in the missing data literature. Two assumptions are built into this notation and deserve naming. The first is that the treatment received by one unit does not affect the outcomes of others, so that Yᵢ depends only on Dᵢ and not on Dⱼ. The second is that there is only one version of treatment, so that "treated" means the same thing for everyone. Together these are known as the stable unit treatment value assumption. It fails in obvious ways in vaccination programmes, where immunising your neighbour protects you, and in marketing experiments where treated customers talk to untreated ones. It fails in subtler ways when "the treatment" bundles together interventions that differ across sites. The methods in this book take it as given; when it is doubtful, the estimand itself must be redefined before any estimation begins. Choosing the estimand Because individual effects are unobservable, causal analyses target summaries of them. The choice of summary is not a technicality. It determines what question the analysis answers, and different summaries can have different signs in the same data. The average treatment effect, E[Y(1) − Y(0)], is the mean effect across the whole population. It answers the question of what would happen, on average, if everyone were treated rather than no one. The average treatment effect on the treated, E[Y(1) − Y(0) | D = 1], restricts attention to those who actually received treatment. It is often the more natural target when evaluating an existing programme, because the policy question is whether the programme helped the people it reached, not whether it would have helped people it was never designed for. The conditional average treatment effect, E[Y(1) − Y(0) | X = x], describes how the average effect varies with observed characteristics X. It is the building block for personalisation: for deciding who should be treated. Beyond these, there are estimands defined by instruments, such as the local average treatment effect of Guido Imbens and Joshua Angrist, which describes the effect among "compliers" whose treatment status is moved by an instrument; estimands defined by policies, such as the expected outcome if treatment were assigned according to some rule; and estimands defined by continuous or multi-valued treatments, such as a dose-response curve. The ones used most often in this book, and the questions each answers, are set out in Table 1. Table 1. Common causal estimands and the questions they answer. Estimand Definition Question answered Typical use ATE E[Y(1) − Y(0)] Effect of treating everyone versus no one Population-wide policy ATT E[Y(1) − Y(0) given D = 1] Effect on those actually treated Programme evaluation CATE E[Y(1) − Y(0) given X = x] Effect for units with characteristics x Targeting, personalisation LATE Effect among compliers with an instrument Effect for those moved by the instrument Instrumental variables designs Policy value E[Y(π(X))] for a rule π Mean outcome if treatment follows rule π Choosing whom to treat The practical lesson is that the estimand must be named before the method is chosen, and named in terms of the decision the analysis is supposed to inform. A pharmaceutical regulator deciding whether to approve a drug for a population needs something like an average effect in that population. A clinician deciding whether to prescribe it to the patient in front of her needs something closer to a conditional effect given that patient's characteristics. A health system deciding how to allocate a limited number of places in a programme needs a policy value, or at least a ranking of patients by expected benefit. These are three different targets, and a single analysis that reports "the effect" without specifying which one it means is not finished. The estimand is also where the most consequential choices about the target population are made. The ATE in a sample of trial volunteers is not the ATE in the population of patients who will eventually receive a drug, and no amount of statistical sophistication within the sample can close that gap without additional assumptions about how the two populations differ. Machine learning has made it easier to estimate effects conditional on many covariates, and therefore, in principle, easier to transport them to new populations with different covariate distributions. But transport requires the assumption that the conditional effects themselves are the same across populations, which is a claim about the world, not about the model. Identification: from potential outcomes to data An estimand is defined in terms of potential outcomes, which we cannot observe. Identification is the step that expresses it in terms of the distribution of variables we can observe. This step is where causal knowledge enters the analysis, and it is logically prior to estimation. If an estimand is identified, then with enough data it could be computed exactly; the remaining questions are statistical. If it is not identified, then even an infinite sample would not reveal it, and no estimator, however clever, can help. In a randomised experiment, identification is straightforward. Because treatment is assigned by a coin flip, it is independent of the potential outcomes, and the difference in mean observed outcomes between treated and control groups equals the ATE. This is why randomised trials occupy their privileged place: the design itself delivers identification without any assumption about how covariates relate to outcomes. In observational data, the standard route to identification rests on three assumptions. The first is conditional exchangeability, also called unconfoundedness, ignorability or selection on observables depending on the discipline: within groups of units that share the same covariates X, treatment is as good as randomly assigned. Formally, the potential outcomes are independent of D given X. The second is overlap or positivity: every unit has a probability of treatment strictly between zero and one given its covariates, so that for every kind of treated unit there are comparable untreated units, and vice versa. The third is consistency, which links the observed outcome to the potential outcome under the treatment actually received and is closely related to the no-multiple-versions condition above. Under these assumptions, the ATE can be written in two ways that will recur throughout the book. The first uses the outcome regressions μ₁(x) = E[Y | D = 1, X = x] and μ₀(x) = E[Y | D = 0, X = x]: the ATE equals the average over the population of μ₁(X) − μ₀(X). This is sometimes called the g-formula or standardisation, and it says that we can compare treated and untreated units within each covariate stratum and then average across strata. The second uses the propensity score e(x) = P(D = 1 | X = x), introduced by Paul Rosenbaum and Donald Rubin in 1983: the ATE equals the average of D·Y/e(X) − (1 − D)·Y/(1 − e(X)). This is inverse probability weighting, and it says that we can reweight the observed sample so that treated and untreated groups each resemble the whole population. Both formulas are exact under the identifying assumptions. Both involve functions, μ and e, that must be estimated from data. And both, as the next two chapters show, behave very differently when those functions are estimated by flexible machine learning methods. But the point to register here is that neither formula comes from the data. The claim that conditioning on X removes confounding is a substantive assertion about the world, justified, if at all, by knowledge of how treatment decisions were made. What machine learning cannot supply It is tempting to think that the more covariates one includes, the more plausible unconfoundedness becomes, and that machine learning, by making it practical to condition on hundreds or thousands of variables, therefore strengthens identification. This is only partly true, and the part that is false causes real damage. It is true that many observational analyses have been undermined by conditioning on too few variables, often because the linear models used could not accommodate more without overfitting. Machine learning relaxes that constraint. If the relevant confounders are among the measured variables but interact with one another in complicated ways, a flexible model can adjust for them where a linear regression with main effects could not. But adding variables can also make things worse. Conditioning on a variable that is affected by the treatment, a mediator or a later consequence, blocks part of the effect one is trying to estimate or opens spurious paths. Conditioning on a collider, a variable caused by both treatment and outcome or by their causes, can create association where there was none. Conditioning on a strong predictor of treatment that has no relationship with the outcome, an instrument in disguise, does nothing to reduce confounding bias but inflates variance and, when some unmeasured confounding remains, can amplify the bias that is left. The graphical framework developed by Judea Pearl and others gives precise rules, the back-door criterion among them, for deciding which sets of variables are sufficient for adjustment, and those rules depend on the causal structure, not on how well the variables predict anything. This is why "throw every available variable into the model and let the algorithm sort it out" is not a causal strategy. The algorithm can sort variables by their predictive value for outcome or treatment. It cannot sort them by their causal role, because causal role is not a property of the joint distribution of the observed data. Two data-generating processes with identical observational distributions can have different causal structures and different treatment effects. The choice of adjustment set therefore has to be made, at least in outline, before the machine learning begins, using subject-matter knowledge about timing and mechanism. A useful working rule is to include pre-treatment variables that plausibly affect the outcome, to exclude anything measured after treatment was determined, and to think carefully before including variables that affect treatment but plausibly not the outcome. The same caution applies to overlap. Machine learning can estimate propensity scores for rich covariate sets, and those estimates frequently reveal that some kinds of units are almost always treated or almost never treated. That is useful information: it shows where the data contain no comparison. But it is information about a limitation, not a solution to one. When overlap fails, the average treatment effect in the full population is not identified from the data without extrapolation, and a flexible model's extrapolation is no more trustworthy than a linear model's. Chapter 8 returns to this point, because the high-dimensional settings in which machine learning is most attractive are exactly those in which strict overlap is least likely to hold. Beyond selection on observables Unconfoundedness is a strong assumption, and in many applications it is not credible. The econometric tradition has developed designs that achieve identification by other means. An instrumental variable is a source of variation in treatment that affects the outcome only through treatment and is itself as good as random: lotteries for school places, distance to a hospital, the random assignment of judges of differing severity. Regression discontinuity exploits rules that assign treatment on the basis of a threshold in a running variable. Difference-in-differences compares changes over time between groups that were and were not exposed to a policy, assuming their trends would otherwise have been parallel. Each design replaces unconfoundedness with a different assumption, and each has been extended with machine learning in the same spirit as the methods of this book: the design provides identification, and flexible learners estimate the nuisance functions that the identified estimand depends on. Double machine learning, in particular, handles instrumental variable models naturally, and Chapter 4 describes how. The general lesson carries over. Machine learning enters after identification has been secured and serves to estimate, as flexibly as possible, the quantities that identification has shown to be relevant. Why the distinction matters in practice A reader who has absorbed the argument so far might wonder why causal machine learning needs a separate literature at all. If identification is settled first and machine learning simply estimates the regression μ or the propensity score e, why not fit the best predictive model available for each and plug it into the relevant formula? The answer, which occupies the next chapter, is that a good predictor of Y or of D is not automatically a good ingredient for estimating the treatment effect. Predictive models are tuned to minimise prediction error, and the errors they are allowed to make are chosen with that goal in mind. Those errors do not cancel when the model is plugged into a causal formula; they propagate into the estimate, often in systematic directions and at rates that swamp the sampling uncertainty. The resulting bias can be large even when the model predicts beautifully, and, more insidiously, it invalidates the standard errors, so that the analyst is not only wrong but confidently wrong. Seeing why this happens is the key to understanding every method in the rest of the book. Double machine learning, targeted learning and generalized random forests are, at bottom, three answers to one question: how can we use predictive models that are good but inevitably biased as components of an estimator that is itself unbiased enough, at the scale of its own standard error, to support inference? The question makes sense only once the estimand and its identification are fixed. That is why they come first. Chapter 2. Why Plugging In Prediction Fails The simplest way to combine machine learning with causal inference is to take one of the identification formulas from the previous chapter, estimate the functions it requires with the best available predictive model, and compute the result. Fit a gradient-boosted model for the outcome given treatment and covariates, predict each unit's outcome under treatment and under control, and average the difference. Or fit a random forest for the probability of treatment, and use its predictions as weights. These plug-in estimators are intuitive, easy to implement and, in general, wrong in a way that matters. This chapter explains why, because the explanation contains the whole logic of the remedy. A model that economists already use The clearest way to see the problem is through a model that has been the backbone of applied econometrics for decades: the partially linear regression. It supposes that the outcome depends on a treatment D with a constant effect θ, plus an unknown function g of covariates X, plus noise: Y = θ·D + g(X) + U, with E[U | X, D] = 0. Treatment itself depends on the covariates through another unknown function m: D = m(X) + V, with E[V | X] = 0. In words, the covariates confound the relationship between D and Y, because they affect both, and we want θ, the effect of D on Y holding X fixed. Traditional practice replaces g with a linear function of X, perhaps with a few interactions and squares chosen by the analyst, and runs ordinary least squares. When X contains many variables or g is highly nonlinear, this is a serious restriction: if the linear specification misses part of g that is correlated with D, the estimate of θ absorbs the omitted confounding and is biased. The motivation for machine learning is exactly to estimate g without committing to a functional form. The natural first attempt is iterative. Guess θ, estimate g by fitting a random forest or lasso to Y − θ·D on X, update θ by regressing Y − ĝ(X) on D, and repeat until the estimates settle. Or, more directly, split the sample, estimate g on one half and regress Y − ĝ(X) on D in the other. Either way the final step is a regression of the residualised outcome on the treatment. It looks like a sensible generalisation of what econometricians have always done. Regularisation bias The trouble lies in what the machine learning model does to g. Every modern predictive method controls its variance by some form of regularisation: the lasso shrinks coefficients toward zero, random forests average over trees grown on subsamples and stop splitting when leaves become small, boosting takes small steps and stops early, neural networks use weight decay and dropout. Regularisation deliberately introduces bias in exchange for lower variance, and in high-dimensional problems this trade is what makes prediction possible at all. For estimating g, the resulting ĝ converges to the truth, but more slowly than the familiar root-n rate of parametric statistics. Typical rates for nonparametric learners in reasonably high dimensions are closer to n to the power minus one quarter, or slower. Now examine what this slow convergence does to the estimate of θ. Writing out the error of the naive estimator, as Chernozhukov and co-authors do in their 2018 paper introducing double/debiased machine learning, it separates into two parts. One part is a well-behaved average of noise terms, which, scaled by root-n, converges to a normal distribution just as in textbook regression. The other part is driven by the product of the treatment's dependence on X and the error in ĝ: roughly, the scaled sum over observations of m(Xᵢ) multiplied by the gap between the true g(Xᵢ) and its estimate. Because ĝ converges more slowly than root-n, this second term does not vanish when multiplied by root-n. It grows. The estimator is consistent, in the sense that it eventually gets close to θ, but its bias shrinks more slowly than its standard error. At any realistic sample size the confidence interval is centred in the wrong place, and as the sample grows the problem, measured in units of standard error, gets worse rather than better. The mechanism is intuitive once stated. Regularisation shrinks ĝ toward something simpler than the true g. Whatever part of g the learner fails to capture is left in the residual Y − ĝ(X). Since g and m are both functions of the same covariates, the uncaptured part of g is typically correlated with D through m. The final regression therefore attributes some of the omitted confounding to the treatment, exactly as an under-specified linear regression would. Machine learning has not removed omitted variable bias; it has reduced it to a level that is small in absolute terms but still large relative to the precision the analyst claims. Simulations make the damage vivid. The 2018 paper shows that, with a moderately complex g and a random forest used to estimate it, the naive estimator's sampling distribution is visibly shifted away from the true θ, while the corrected estimator described in the next chapter is centred on it. The shift is not small noise. It is a systematic displacement of a size comparable to the width of the distribution, which means that nominal ninety-five per cent confidence intervals cover the truth far less often than advertised. An earlier warning from the lasso The same lesson had already been learned in a narrower setting. In the early 2010s, economists began using the lasso to select control variables from large candidate sets. The obvious procedure was to regress the outcome on the treatment and all candidate controls with a lasso penalty on the controls, keep the controls with non-zero coefficients, and then run ordinary least squares of the outcome on the treatment and the selected controls. Alexandre Belloni, Victor Chernozhukov and Christian Hansen showed in 2014 that this procedure can fail badly. The lasso selects variables that predict the outcome well. A confounder that strongly affects treatment but only moderately affects the outcome may be dropped, because its predictive contribution to Y is small relative to the penalty. Yet omitting it biases the treatment effect precisely because it is strongly related to D. Worse, whether it is dropped depends on the particular sample, so the post-selection estimator has a distribution that is not normal and not centred at the truth, and standard confidence intervals are unreliable. This is a specific instance of a general result in the post-selection inference literature, associated particularly with Hannes Leeb and Benedikt Pötscher, that inference after data-driven model selection is fragile. Their remedy was double selection: run one lasso of Y on the controls and a second lasso of D on the controls, and then include the union of the variables selected by either in the final regression. A confounder that matters strongly for treatment will be picked up by the second lasso even if the first misses it. The resulting estimator is robust to the selection mistakes that each lasso inevitably makes, and it supports valid inference under conditions the authors spell out. The name "double" in double machine learning descends from this idea: learn both the outcome's and the treatment's dependence on covariates, and use both. Overfitting bias Regularisation bias is the first of two problems. The second arises when the same data are used both to fit the nuisance functions and to estimate the treatment effect. Flexible learners can fit the training data very closely. A deep random forest or a large boosted ensemble will, on the observations it was trained on, produce predictions that track the noise as well as the signal. If the residuals Y − ĝ(X) are computed on those same observations, they are artificially small and systematically related to the noise in the outcome. Using them in the final regression introduces a bias term involving the correlation between the estimation error in ĝ and the outcome noise, which need not vanish at the required rate. Classical semiparametric theory avoided this problem by imposing conditions, known as Donsker conditions, that limit how complex the class of estimated functions can be. Those conditions are satisfied by many traditional smoothers but are hard to verify, and often false, for the adaptive, high-capacity learners that practitioners actually use. The remedy for overfitting bias is sample splitting: estimate the nuisance functions on one part of the data and evaluate them on another. Because the evaluation sample was not used for fitting, the errors in the nuisance estimates are independent of the noise in the evaluation observations, and the problematic term averages out. Splitting on its own wastes half the data, which is why the methods of this book use cross-fitting, rotating the roles of the folds so that every observation is used for estimation and every observation also receives nuisance predictions from a model that did not see it. Cross-fitting is described in detail in Chapter 4. Here it is enough to note that sample splitting fixes overfitting bias but does nothing about regularisation bias. A naive estimator with sample splitting is still biased, because the omitted part of g remains correlated with D whether or not ĝ was fitted on the same observations. The same problem, other formulas The partially linear model makes the mechanism easy to see, but the problem is general. Consider the two identification formulas for the ATE from Chapter 1. The outcome regression estimator averages μ̂₁(X) − μ̂₀(X), where μ̂₁ and μ̂₀ are machine learning models of the outcome in the treated and control arms. Its error is essentially the average of the errors in μ̂₁ and μ̂₀ across the population. Those errors are dominated by regularisation bias in regions where data are sparse, and they are systematic: a learner that shrinks toward the overall mean will understate differences between arms in regions where treated units are rare. The error shrinks at whatever rate the learner converges, which is typically slower than root-n, so the estimator again has bias that dominates its standard error. There is also no general way to compute a valid standard error for it, since the learner's own sampling behaviour is complicated and poorly characterised. The inverse probability weighting estimator divides by ê(X), a machine learning estimate of the propensity score. Its bias depends on the error in ê, again typically converging slowly, and the division makes matters worse: where the true propensity is small, small absolute errors in ê produce large relative errors in the weights. Machine learning classifiers are also frequently poorly calibrated, meaning that their predicted probabilities do not match observed frequencies, particularly near zero and one. A boosted model tuned for classification accuracy can push predicted probabilities toward the extremes, producing enormous weights for a handful of observations and an estimator whose variance is dominated by them. Neither formula, used alone with a machine learning estimate plugged in, delivers what the analyst needs: an estimate whose error is dominated by sampling noise of a known, root-n size, so that a normal approximation and a confidence interval are available. Good prediction, wrong target A deeper point sits underneath both problems. When a learner is tuned by cross-validation to predict Y as accurately as possible, it is optimising a criterion that weighs all errors in the prediction of Y equally, wherever in covariate space they occur and whatever their relation to treatment. For the causal estimate, however, errors are not equally costly. An error in ĝ that is uncorrelated with D does no harm to the estimate of θ; it only adds noise. An error that is correlated with D, even a small one, translates directly into bias. The predictive criterion is blind to this distinction, so the learner happily trades away accuracy in exactly the directions that matter most for the causal question if doing so buys accuracy elsewhere. Consider a stylised evaluation of a job-training programme. Suppose participation is strongly driven by recent unemployment spells, which also depress future earnings, while future earnings are driven mainly by education and prior earnings, with recent unemployment playing a smaller role. A learner fitted to predict future earnings will concentrate its capacity on education and prior earnings, where most of the predictable variation lies, and will model the unemployment effect coarsely, perhaps shrinking it heavily. From a predictive standpoint this is the right choice: the unemployment variable adds little to the accuracy of earnings forecasts. From a causal standpoint it is the worst possible choice, because the part of the earnings process the learner neglected is precisely the part that is entangled with participation. The residual earnings after subtracting the learner's predictions will still contain a component that is lower among participants, and the programme will look less effective than it is. This is the same mechanism as the dropped confounder in the naive lasso, generalised to any learner. It explains a pattern that practitioners sometimes find baffling: switching to a more accurate outcome model, as judged by held-out error, does not reliably improve a plug-in treatment effect estimate and can make it worse. Predictive accuracy on Y is simply not the quantity that governs the causal error. What governs it is how well the learner captures the part of the outcome process that co-varies with treatment, and no predictive criterion measures that directly. The remedy must therefore change the structure of the estimator rather than rely on ever-better predictions. What an answer must look like It helps to be precise about the requirement. Causal inference with flexible nuisance estimation asks for an estimator θ̂ such that root-n times (θ̂ − θ) is approximately normal with mean zero and a variance that can be estimated. The nuisance functions will be estimated at slower rates; that is unavoidable once we refuse to assume a parametric form. So we need the error in the final estimate to depend on the nuisance errors only weakly, weakly enough that slow convergence of the nuisances does not contaminate the fast convergence of θ̂. One way to achieve this would be to find estimators whose error depends on the nuisance errors only through a product of two of them. If the error in the outcome model is of order n to the minus one quarter, and the error in the propensity model is also of order n to the minus one quarter, their product is of order n to the minus one half, which is exactly the scale of sampling noise. Multiplied by root-n, it vanishes. Each nuisance can then be learned at a pedestrian rate, one achievable by lasso, random forests, boosting or neural networks under reasonable conditions, while the combined estimator converges at the parametric rate. Estimators with this property exist, and they are not new. They are built from what semiparametric statisticians call influence functions and what econometricians, following Neyman, call orthogonal scores. The augmented inverse probability weighted estimator of James Robins, Andrea Rotnitzky and Lue Ping Zhao, published in 1994, has it. So does the partialling-out estimator for the partially linear model that Peter Robinson analysed in 1988, which regresses the residual Y − E[Y | X] on the residual D − E[D | X]. The contribution of the modern literature was to see that these old constructions are precisely what is needed to make machine learning usable for causal inference, to state the conditions under which they work with arbitrary learners, and to combine them with cross-fitting so that the conditions become easy to meet. It is worth pausing on Robinson's estimator, because it is the simplest example of the remedy and the direct ancestor of both double machine learning and the causal forest. Instead of residualising only the outcome, it residualises both outcome and treatment on the covariates, and then regresses one residual on the other. If the outcome model ℓ(X) = E[Y | X] is slightly wrong and the treatment model m(X) = E[D | X] is slightly wrong, the errors enter the final regression only through their product. The treatment residual D − m̂(X) is approximately uncorrelated with any function of X, including the part of the outcome that the outcome model missed, so the omitted confounding no longer leaks into the estimate. This is the Frisch–Waugh–Lovell theorem of linear regression, made robust to the imperfect nonparametric estimation of both projections. The next chapter explains why it works, and why the same principle generalises to almost any estimand one might care about. Chapter 3. Orthogonal Scores and the Influence Function Every method in this book depends on one idea, which appears under several names. Econometricians speak of Neyman orthogonality and orthogonal moment conditions. Biostatisticians speak of efficient influence functions, efficient influence curves and doubly robust estimating equations. Machine learning researchers increasingly speak of debiasing. This chapter explains the idea without assuming a background in semiparametric theory, because once it is understood the three families of methods become variations on a theme rather than separate techniques to be memorised. Estimating equations and their sensitivity Most estimators can be written as the solution to an estimating equation: choose θ so that the sample average of some score function equals zero. Ordinary least squares, for instance, sets the average of the residual times each regressor to zero. Maximum likelihood sets the average score, the derivative of the log-likelihood, to zero. In causal problems, the score typically depends not only on the parameter of interest θ but also on nuisance functions such as the outcome regression and the propensity score. Write the nuisances collectively as η. The estimator solves average of ψ(Wᵢ; θ, η̂) = 0, where Wᵢ collects the observed data for unit i and η̂ is the estimated nuisance. The question the previous chapter raised is how sensitive θ̂ is to errors in η̂. Suppose we perturb the nuisance slightly from its true value, moving η in some direction by a small amount r. The expected score changes. If it changes in proportion to r, then an error of size r in the nuisance produces an error of roughly size r in θ̂. That is the situation of the naive plug-in estimators: nuisance errors pass through one-for-one, and slow nuisance convergence means slow convergence of θ̂. Neyman orthogonality is the requirement that the derivative of the expected score with respect to the nuisance, evaluated at the truth, be zero in every direction. If this holds, then a perturbation of size r in the nuisance changes the expected score only by an amount proportional to r squared, or to a product of errors in different nuisance components. Small errors become very small errors. An estimator built on such a score inherits this insensitivity: to first order, it behaves as though the true nuisance functions were known. Jerzy Neyman introduced the idea in 1959 in the context of hypothesis testing with nuisance parameters, as the basis of his C(α) tests. He noticed that a test statistic could be made locally insensitive to the estimation of nuisance parameters by projecting the score for the parameter of interest onto the space orthogonal to the scores for the nuisances. That projection is what gives the idea its name. The modern contribution is to apply it when the nuisance is not a handful of parameters but an entire function learned by a flexible algorithm. The partialling-out score The partially linear model from the previous chapter provides the simplest example. The naive score, which regresses Y − g(X) on D, is ψ = (Y − θ·D − g(X))·D. Differentiating its expectation with respect to g in any direction gives a term proportional to the expected value of D times the direction, which is not zero because D depends on X. The score is sensitive to errors in g. The orthogonal score residualises both variables. Let ℓ(X) = E[Y | X] and m(X) = E[D | X]. The score is ψ = (Y − ℓ(X) − θ·(D − m(X)))·(D − m(X)). Setting its sample average to zero gives Robinson's estimator: the regression of the outcome residual on the treatment residual. Differentiate with respect to ℓ: the result involves E[D − m(X) | X], which is zero by the definition of m. Differentiate with respect to m: the result involves E[Y − ℓ(X) − θ·(D − m(X)) | X] and a term in E[D − m(X) | X], both of which vanish at the truth. The score is orthogonal to both nuisances. The intuition is worth making explicit. The treatment residual D − m(X) is the part of treatment that cannot be predicted from covariates, the "as good as random" variation that identification rests on. Because it is by construction uncorrelated with every function of X, any error in the outcome model, which is a function of X, is also uncorrelated with it. The error washes out in the final regression instead of being attributed to treatment. Symmetrically, errors in m shift the treatment residual by a function of X, but the outcome residual is also uncorrelated with functions of X at the truth, so these errors wash out too. Only when both models are wrong in related ways does an error survive, and then it enters as a product. Influence functions The partialling-out score was easy to find because the model is simple. For other estimands, a systematic route to orthogonal scores runs through the influence function, a concept from robust statistics and semiparametric theory. The influence function of an estimand describes how the estimand changes when the distribution of the data is perturbed slightly toward a point mass at a single observation. Informally, it measures each observation's contribution to the estimate. For well-behaved estimators, the estimation error is approximately the sample average of the influence function evaluated at each observation, which is why the influence function determines the estimator's asymptotic variance: the variance of θ̂ is approximately the variance of the influence function divided by n. For an estimand such as the ATE, many estimators are possible, each with its own influence function. Among them there is one with the smallest variance, the efficient influence function. Its variance defines the semiparametric efficiency bound, the lowest variance any regular estimator can achieve without additional assumptions. For the ATE under unconfoundedness, this bound was derived by Jinyong Hahn in 1998, building on earlier work by James Robins and colleagues. The general theory is set out in the 1993 monograph by Peter Bickel, Chris Klaassen, Ya'acov Ritov and Jon Wellner. The crucial property for our purposes is that the efficient influence function, viewed as a score, is automatically Neyman orthogonal. This is not a coincidence. The efficient influence function is constructed by projecting out every direction in which the nuisance could vary, which is exactly what orthogonality requires. So a practical recipe for building a debiased estimator for any smooth estimand is to find its efficient influence function and use it as the estimating equation. Oliver Hines, Oliver Dukes, Karla Diaz-Ordaz and Stijn Vansteelandt give an accessible account of how to do this in a 2022 article in The American Statistician, and Aaron Fisher and Edward Kennedy offer a visual introduction to the same ideas in a 2021 article in the same journal. The doubly robust score for the average treatment effect The efficient influence function for the ATE under unconfoundedness produces the estimator that sits at the centre of modern practice. With μ₁(x) and μ₀(x) the outcome regressions in the treated and control arms and e(x) the propensity score, the score for each observation is μ₁(X) − μ₀(X) + D·(Y − μ₁(X))/e(X) − (1 − D)·(Y − μ₀(X))/(1 − e(X)) − θ. The first part, μ₁(X) − μ₀(X), is the outcome-regression estimate of the individual contrast. The second part is a correction: the inverse-probability-weighted average of the outcome model's residuals. If the outcome model were perfect, the residuals would average to zero within each covariate stratum and the correction would vanish. If the outcome model is biased in some region, the correction measures and removes that bias using the observed outcomes, weighting by the inverse propensity so that the removal is representative of the whole population. Averaging this score over the sample gives the augmented inverse probability weighting estimator, usually abbreviated AIPW, introduced by Robins, Rotnitzky and Zhao in 1994. It is also what Heejung Bang and James Robins, in an influential 2005 paper, described as doubly robust, and the term has stuck. The estimator is consistent if either the outcome model or the propensity model is correctly specified, even if the other is wrong. If the outcome model is right, the correction term has mean zero regardless of the propensity. If the propensity model is right, the weighting correction exactly offsets any error in the outcome model on average. Double robustness was originally valued as insurance against misspecification of parametric models: an analyst fitting logistic and linear regressions could be wrong about one and still obtain a consistent estimate. In the machine learning setting, its more important consequence is quantitative. The bias of the AIPW estimator, when both nuisances are estimated with error, is governed by the product of the two errors. Specifically, the leading bias term is proportional to the expected product of the propensity error and the outcome-regression error, roughly the integral of (e − ê)·(μ − μ̂) weighted by the inverse of the estimated propensity. This is the mixed bias, or product rate, property. The product rate condition The product structure is what makes machine learning usable. For the bias to be negligible relative to the standard error, which is of order one over root-n, the product of the two nuisance errors must shrink faster than one over root-n. This holds, for instance, if both errors shrink faster than n to the minus one quarter. It also holds if one nuisance is estimated very well and the other rather poorly: a nearly known propensity score, as in a randomised experiment with known assignment probabilities, compensates for a crude outcome model, and vice versa. The rate n to the minus one quarter is attainable by many machine learning methods under reasonable conditions. The lasso achieves it when the true regression is approximately sparse, meaning that a modest number of variables capture most of the signal. Random forests, boosted trees and neural networks achieve it under various smoothness or structural assumptions that have been established in the theoretical literature over the past decade. None of these guarantees is unconditional, and in any particular application one cannot verify that the rate holds. But the condition is far weaker than the root-n rate that plug-in estimators would require, and it is weaker than the correct specification that parametric methods require. That relaxation, from "the model must be right" to "the models must be reasonably good, jointly", is the practical gain. It is worth being clear about what the rate condition does not say. It does not say that any learner will do, or that nuisance quality does not matter. The product must be small, and in finite samples the constant in front of it matters as much as the rate. Poorly tuned learners, or learners that systematically miss the structure of the problem, produce product terms that are not negligible at realistic sample sizes. Chapter 4 discusses how to choose and tune learners with this in mind. The limits of double robustness Double robustness is sometimes presented as if it made an estimator safe. The history of the idea contains a useful corrective. In 2007, Joseph Kang and Joseph Schafer published a simulation study in Statistical Science, pointedly titled "Demystifying double robustness", in which both the outcome and propensity models were moderately misspecified in a realistic way. Several doubly robust estimators performed worse than a simple outcome regression, some of them dramatically so. The reason was that the misspecified propensity model produced a few very small estimated probabilities, and the inverse weights attached to those observations were enormous. The correction term, which is supposed to repair the outcome model's errors, instead amplified noise from a handful of units. The discussion that followed, with responses from Robins, Rotnitzky, van der Laan and others, clarified several points that remain important. First, double robustness is a statement about consistency, about what happens as the sample grows; it offers no guarantee about finite-sample performance when both models are wrong. Second, the variance of inverse-weighted corrections depends heavily on overlap, and estimators that let weights explode inherit that instability. Third, there are ways of constructing doubly robust estimators that are far more stable, for example by keeping predictions within the range of the observed outcomes, which is one of the motivations for the targeted learning approach of Chapter 5. In the machine learning setting, the lesson translates directly. Flexible propensity models can produce predicted probabilities very close to zero or one, especially when many covariates are available and the model is allowed to overfit. A doubly robust estimator built on such predictions is formally orthogonal and practically fragile. The product rate condition assumes, among other things, that the propensity is bounded away from zero and one; when it is not, the constant multiplying the product of errors grows without limit and the theoretical guarantees become empty. Diagnosing and responding to weak overlap is therefore not an optional refinement but part of the method, and Chapter 8 treats it in detail. Why orthogonality alone is not enough Orthogonality addresses regularisation bias. It does not by itself address overfitting bias, the second problem from the previous chapter. The proof that an orthogonal-score estimator behaves as if the nuisances were known requires controlling a term involving the empirical process: the difference between the sample average and the population average of the score, evaluated at the estimated rather than the true nuisance. Classical theory handles this term by restricting the complexity of the nuisance estimators, the Donsker conditions mentioned earlier. Modern learners often violate those conditions. Cross-fitting removes the need for them. If the nuisance used to evaluate the score for observation i was fitted without observation i, then conditional on the fitted nuisance the score terms are independent across the evaluation fold, and the troublesome empirical-process term can be bounded with elementary arguments that require only that the nuisance errors shrink. This combination, an orthogonal score plus cross-fitting, is the complete recipe of double/debiased machine learning, and its essential logic is shared by every other method in this book. Standard errors for free A further benefit of building estimators on influence functions is that inference follows almost automatically. Because the estimation error is approximately the average of the influence function across observations, the variance of the estimator can be estimated by the sample variance of the estimated score values divided by n. There is no need to bootstrap the whole machine learning pipeline, which would be computationally expensive and, for some learners, theoretically unjustified. The analyst computes each observation's score contribution, takes their mean for the point estimate and their standard deviation for the standard error, and forms a confidence interval in the usual way. This simplicity is easy to underrate. Earlier attempts to use flexible models for causal effects often foundered on inference: an estimate from a random forest or a neural network could be computed, but no honest measure of its uncertainty was available. The influence function approach converts the question of inference about a complex pipeline into the question of estimating the variance of a scalar quantity, which is the most elementary task in statistics. Beyond the average treatment effect The same construction extends well beyond the ATE. The average treatment effect on the treated has its own efficient influence function, involving the outcome model in the control arm and the odds of treatment. Instrumental variable estimands, such as the local average treatment effect, have orthogonal scores built from models of the outcome, treatment and instrument given covariates. Effects of continuous treatments, mediation effects, and the value of treatment policies all admit orthogonal scores, as do the parameters of the best linear approximation to a heterogeneous effect. Victor Chernozhukov, Whitney Newey and Rahul Singh have shown, in work published in Econometrica in 2022, that for a broad class of estimands that are linear functionals of a regression, the correction term can be learned automatically by estimating a function called the Riesz representer, without deriving its analytic form. For the ATE, the Riesz representer is the familiar inverse-propensity weight; for other estimands it has no simple closed form, and learning it directly avoids the numerically fragile step of inverting estimated probabilities. Edward Kennedy's 2022 review, "Semiparametric doubly robust targeted double machine learning", whose title deliberately strings together the names of the competing traditions, argues that these literatures describe one body of theory. Having reached this point, a reader can see why. Double machine learning takes the efficient influence function, uses it as an estimating equation, and adds cross-fitting. Targeted maximum likelihood, described in Chapter 5, takes the same efficient influence function and uses it to adjust an initial estimate so that the equation is solved by a substitution estimator. The causal forest of Chapter 7 uses the partialling-out score locally, within neighbourhoods defined by a forest, to estimate how the effect varies. The practical differences between these methods are real and worth understanding. But they sit on top of a common foundation, and it is the foundation, more than any particular learner, that makes their results trustworthy. Hashtags: #CausalMachineLearning #CausalInference #TreatmentEffectEstimation #PotentialOutcomes #AverageTreatmentEffect #ConditionalAverageTreatmentEffect #HeterogeneousTreatmentEffects #DoubleMachineLearning #DebiasedMachineLearning #TargetedMaximumLikelihood #TMLE #GeneralizedRandomForests #CausalForests #NeymanOrthogonality #InfluenceFunctions #DoubleRobustness #CrossFitting #PropensityScores #ConfoundingAdjustment #NuisanceFunctions #PolicyLearning #MetaLearners #TreatmentEffectHeterogeneity #SemiparametricInference #FutureOfCausalMachineLearning
- Advanced Microscopy (Confocal, Two-Photon, and Super-Resolution Techniques)
Download the Book (PDF): Introduction A fluorescent molecule is a small and unreliable lamp. Excite it, and it will emit a photon a few nanoseconds later, most of the time. Excite it again and it will usually oblige again. But on each cycle there is a small chance that something else happens instead: the molecule slips into a long-lived dark state, reacts with oxygen, or is chemically altered so that it will never fluoresce again. A good organic dye might survive on the order of a million excitation cycles before this happens; a fluorescent protein often far fewer. And the objective lens, however expensive, collects only a fraction of the photons that are emitted, and the detector converts only a fraction of those into signal. Every image a fluorescence microscope has ever produced was assembled from this finite and fragile supply. This book is organised around a single consequence of that fact. Resolution, speed, imaging depth and the health of the specimen all draw on the same limited account of photons, and every technique described in the chapters that follow is a different way of spending it. Confocal microscopy spends photons to buy optical sectioning, throwing away most of the light it excites in order to keep the light it wants. Two-photon microscopy spends laser power, and eventually tissue heating, to buy depth. Light-sheet microscopy economises, exciting only the plane it is looking at, and returns the savings as speed and gentleness. Stimulated emission depletion (STED) and single-molecule localization techniques such as PALM and STORM spend photons lavishly, sometimes brutally, to buy resolution far below the classical diffraction limit. Adaptive optics and deconvolution try to recover value from photons that the optics have already squandered through aberration and blur. Seen this way, the question a microscopist faces is rarely "which instrument has the best resolution?" It is "what does my question actually require, and what am I willing to pay for it?" A live zebrafish embryo tracked for two days needs gentleness and speed more than it needs 30-nanometre resolution. A fixed synapse whose protein arrangement is the whole point of the experiment can afford an extravagant photon budget, provided the labels and the analysis are sound. A neuron firing half a millimetre beneath the surface of a mouse cortex needs penetration above almost everything else. The instruments differ, but the reasoning that selects among them is the same, and it is the reasoning, more than the instruments, that this book tries to teach. Why the budget framing matters It is easy to learn advanced microscopy as a catalogue. Each technique has its founding paper, its characteristic resolution figure, its commercial implementations and its known pitfalls, and a reader can memorise these as separate facts. That approach fails in practice, for two reasons. The first is that the headline numbers mislead. A STED microscope may be specified to reach 30 nanometres, but it reaches that figure only with a robust dye, high depletion power and a sample that can tolerate both; on a live cell expressing a dim fluorescent protein the practical figure may be several times worse. A localization microscope may report a precision of 10 nanometres per molecule, but the resolution of the final image also depends on how densely the structure was labelled, how large the antibodies were, and how well drift was corrected. Resolution is not a property of an instrument. It is a property of an instrument, a sample, a label and an acquisition, jointly, and the photon budget is the variable that ties them together. The second reason is that the costs are often invisible. Photobleaching is at least obvious: the image fades. Phototoxicity frequently is not. A cell can look perfectly healthy in the final frame of a time-lapse while having arrested its cell cycle, altered its mitochondrial dynamics or changed the very behaviour under study. Heating from an infrared laser does not announce itself on the monitor. Spherical aberration degrades an image gradually with depth, so that a researcher can mistake an optical artefact for a biological gradient. Deconvolution can produce crisp, convincing structures that are partly the algorithm's invention. A microscopist who thinks in terms of what each photon is buying, and what each exposure is costing, is far better placed to notice these hidden charges than one who thinks in terms of specifications. What this book covers The chapters proceed roughly from the physics that sets the rules, through the major families of instrument, to the corrections that recover what the optics lose. Chapter 1 sets out the diffraction limit and the statistics of photon counting, because these two constraints, one from wave optics and one from quantum noise, define every trade-off that follows. Chapter 2 treats confocal microscopy, the workhorse of biological imaging for four decades, and explains what the pinhole does and does not buy, along with the detector and scanning choices that now matter as much as the pinhole itself. Chapter 3 covers two-photon excitation, the technique that made functional imaging deep in living brains routine, and its three-photon successor. Chapter 4 turns to light-sheet microscopy, whose central idea is almost embarrassingly simple and whose practical consequences for developmental biology and whole-organ imaging have been enormous. Chapter 5 steps back from instruments to the molecules themselves: the photophysics of bleaching and blinking, the chemistry of modern dyes and fluorescent proteins, and the growing evidence on phototoxicity. This chapter sits at the centre of the book because both super-resolution chapters that follow depend on it completely. Chapter 6 explains STED and its relatives, which break the diffraction limit by deliberately switching fluorescence off everywhere except a tiny central region. Chapter 7 covers single-molecule localization microscopy, including PALM, STORM, DNA-PAINT and the hybrid MINFLUX approach, which break the limit by separating molecules in time rather than in space. The final two chapters address the gap between the ideal microscope and the real one. Chapter 8 deals with spherical aberration and other wavefront errors, why they worsen with depth, and how correction collars, matched immersion media and adaptive optics restore performance. Chapter 9 covers deconvolution, from the classical Richardson–Lucy algorithm to learned restoration methods, and argues that it should be treated as a measurement with assumptions rather than as an image-sharpening filter. The conclusion draws these threads into a way of choosing and reporting experiments. It does not summarise the chapters; it argues for a particular discipline of practice that only becomes visible once all of the techniques have been laid side by side. Who this book is for, and what it assumes The intended reader is a scientist or advanced student who already uses a fluorescence microscope, or is about to, and wants to understand the advanced techniques well enough to choose among them, design sound experiments, and read the literature critically. A graduate student in cell biology planning a first super-resolution project, a neuroscientist weighing two-photon against light-sheet approaches, an imaging-facility staff member advising users, and a physicist moving into biological imaging should all find it useful. The book assumes familiarity with the basic components of a widefield fluorescence microscope: objective, filter cube, camera. It uses a small amount of mathematics where the mathematics is the clearest explanation, for example the formula for the diffraction limit or the scaling of localization precision with photon number, but it does not derive results that can be understood without derivation. Where a result is quoted, it is attributed to the work that established it, and the Notes and Further Reading at the end point to the primary papers and to the textbooks that treat each topic in full. Some deliberate omissions should be stated. Electron microscopy, correlative light and electron microscopy, and label-free methods such as coherent Raman scattering and optical coherence tomography are outside the scope of this book, although they appear where they illuminate a comparison. Commercial instruments are mentioned only where a specific product embodies a technique, and never as recommendations; manufacturers' specifications change faster than books do. And although image analysis is a vast and essential subject, it enters here only where it is inseparable from image formation, as it is in localization microscopy and deconvolution. A final remark about tone. Advanced microscopy has a strong tradition of spectacular images, and many of them are genuinely beautiful. The discipline this book argues for is less glamorous: knowing what an image cost, what it can support, and what it cannot. The techniques described here, and the fluorescent labels they depend on, were recognised by the Nobel Prizes in Chemistry of 2008 and 2014, and they have transformed cell biology, and they deserve to be used with the care that their power demands. Chapter 1. The Diffraction Limit and the Photon Budget Two constraints govern everything a fluorescence microscope can do. The first comes from wave optics: light passing through a lens of finite aperture cannot be focused to a point, so every point source in the specimen is imaged as a small blurred spot. The second comes from the quantum nature of light: photons arrive one at a time and at random, so every measurement of brightness carries an irreducible statistical uncertainty. Neither constraint can be engineered away by better manufacturing. Every advanced technique in this book is a strategy for working around one or both of them, and every one of those strategies has a price. Understanding the price requires understanding the constraints first. The point spread function Consider a single fluorescent molecule, far smaller than the wavelength of light, sitting at the focus of an objective lens. It radiates light in all directions. The objective captures the portion that falls within its cone of acceptance and focuses it onto the camera. If lenses were perfect geometric devices, the image would be a point. Instead it is a pattern with a bright central disc surrounded by faint concentric rings, first described for telescopes by George Airy in 1835 and now called the Airy pattern. In three dimensions the pattern is elongated along the optical axis, rather like an hourglass or a pair of cones meeting at their tips, with the brightest region at the focal plane and light spreading out above and below it. This three-dimensional image of a point is the point spread function, or PSF. It is the single most important object in quantitative microscopy. A microscope that images incoherent fluorescence is, to a very good approximation, a linear shift-invariant system: the image of any specimen is the sum of the PSFs of all its fluorescent points, each weighted by its brightness. In mathematical language the image is the convolution of the object with the PSF, plus noise. Almost everything later in this book can be read as an attempt to change the PSF, to measure it, or to undo its effect. The width of the PSF is set by the numerical aperture of the objective, written NA and defined as n sin θ, where n is the refractive index of the medium between lens and specimen and θ is the half-angle of the cone of light the lens accepts. A higher NA means the lens gathers light from a wider range of angles, and it is precisely the high-angle rays that carry information about fine detail. Ernst Abbe, working with Carl Zeiss in Jena in the 1870s, showed that a periodic structure can be resolved only if the lens captures at least the first diffracted orders of light it scatters, and that this sets a minimum resolvable period of roughly λ/(2NA). Lord Rayleigh's criterion, framed for two point sources, gives a closely related figure: two points are just resolved when the centre of one Airy disc falls on the first dark ring of the other, which occurs at a separation of 0.61λ/NA. These formulas differ in their constants because they answer slightly different questions, and practitioners argue about which is most appropriate. The more useful point is that all of them scale in the same way. Resolution improves in proportion to NA and worsens in proportion to wavelength. For green fluorescence at around 520 nanometres and an oil-immersion objective of NA 1.4, the Rayleigh figure is about 230 nanometres. That number, a little under a quarter of a micron, is the diffraction limit in the sense most biologists mean it. Axial resolution is considerably worse. The PSF is typically two to four times longer along the optical axis than it is wide, because the lens captures only a cone of light from one side rather than a full sphere. A commonly used paraxial approximation for the axial extent of a widefield PSF is 2nλ/NA², which reveals the stronger dependence: axial resolution scales with the square of the NA, so it deteriorates rapidly as NA falls. The practical effect is visible in Table 1, which applies the standard lateral and axial formulas to a range of common objectives at an emission wavelength of 520 nanometres. The values are theoretical best cases computed from the formulas, not measured performance, and the axial approximation becomes less accurate at the highest apertures. Table 1. Theoretical lateral (Rayleigh, 0.61λ/NA) and axial (2nλ/NA²) resolution for common objectives at λ = 520 nm. Objective Immersion index n Lateral (nm) Axial (nm) 10×, NA 0.30, air 1.00 about 1,060 about 11,600 20×, NA 0.80, air 1.00 about 400 about 1,600 40×, NA 1.15, water 1.33 about 280 about 1,050 60×, NA 1.30, silicone oil 1.41 about 240 about 870 63×, NA 1.40, oil 1.52 about 230 about 800 100×, NA 1.49, oil 1.52 about 210 about 710 Two lessons follow from the table. First, lateral resolution saturates: going from NA 1.15 to NA 1.49 improves it by only about a quarter. There is a hard ceiling, because NA can never exceed the refractive index of the least refractive medium in the light path, and for biological samples mounted in aqueous media that medium is water, with n of about 1.33. Oil objectives with NA above 1.33 achieve their full aperture only when the specimen itself is mounted in a high-index medium, or when the fluorophore sits within a fraction of a wavelength of the coverslip, as in total internal reflection fluorescence (TIRF) microscopy. Second, axial resolution varies over more than an order of magnitude across the same range, which is why low-NA objectives cannot distinguish one cell layer from the next and why so much of advanced microscopy is concerned with the third dimension. What resolution does and does not mean Resolution in the Rayleigh sense is a statement about separating two equally bright points in the absence of noise. It is not a statement about how precisely a single isolated object can be located, which can be far better, nor about whether an object smaller than the PSF can be detected at all, which it certainly can: a single fluorescent protein a few nanometres across is easily visible if it is bright enough and the background dark enough. It appears as a spot the size of the PSF, but it is visible. Detection, localization and resolution are three different things, and much confusion in the literature arises from treating them as one. This distinction is the seed of an entire family of super-resolution methods. If molecules can be arranged so that only one within any PSF-sized region is emitting at a given moment, each can be located with a precision limited not by diffraction but by the number of photons collected. That idea, developed in Chapter 7, would be useless if photons were free. They are not. Photons are counted, not measured A camera pixel or photomultiplier does not measure light intensity in the way a thermometer measures temperature. It counts discrete photon arrivals, or more precisely the photoelectrons they produce, over an interval. Photon emission from a population of fluorophores is a random process, and the number detected in a fixed interval follows Poisson statistics. The defining property of a Poisson distribution is that its variance equals its mean. If a pixel records on average N photoelectrons, the standard deviation of that count from frame to frame is the square root of N. This gives the fundamental signal-to-noise ratio of any light measurement: N divided by the square root of N, which is simply the square root of N. A pixel that records 100 photoelectrons has a signal-to-noise ratio of 10, meaning its value fluctuates by about ten per cent between identical exposures. To double the signal-to-noise ratio requires four times as many photons. To improve it tenfold requires a hundred times as many. This square-root law, often called shot noise, cannot be removed by any detector, however perfect. It is a property of light itself. Real detectors add their own noise on top. Scientific CMOS cameras, which now dominate widefield and light-sheet imaging, have read noise of around one to two electrons per pixel per frame, plus a small dark current that matters mainly in long exposures. Electron-multiplying CCD cameras can make read noise effectively negligible by amplifying the signal before readout, but the stochastic multiplication process adds an "excess noise factor" of about the square root of two, which is equivalent to halving the quantum efficiency. Photomultiplier tubes and hybrid detectors used in point-scanning systems have their own gain noise and dark counts. The practical consequence is that at very low signal, a few photons per pixel, detector noise can dominate, and choosing the right detector matters enormously. At moderate signal, shot noise dominates and the only way to improve the image is to collect more photons. Quantum efficiency, the fraction of incident photons that produce a detectable photoelectron, varies from around 20 to 45 per cent for conventional photomultipliers, to around 45 per cent or more for gallium arsenide phosphide (GaAsP) photocathodes, to 80 per cent or more for the best back-illuminated cameras in the visible range. These differences translate directly into the photon budget. A detector with twice the quantum efficiency allows the same image quality at half the excitation dose, or twice the imaging duration before bleaching becomes limiting. Where the photons go It is sobering to trace how few of the photons emitted by a sample reach the final image. A fluorophore radiates, roughly speaking, in all directions. An objective of NA 1.4 in oil collects only about a third of the emitted light even in the ideal case; a water-dipping objective of NA 1.0 collects closer to a sixth, and a low-NA objective far less. Each lens element, mirror, dichroic and filter then transmits somewhat less than all of what reaches it; a total transmission of 50 to 70 per cent through the emission path of a complex system is respectable. The detector then converts some fraction of the remainder. Multiply these together and a well-designed widefield system may detect around five to fifteen per cent of emitted photons; a confocal system, with its additional pinhole losses and scanning optics, often detects fewer. Meanwhile each fluorophore can emit only a finite number of photons before it photobleaches, a limit discussed in depth in Chapter 5. The total number of photons a sample can ever yield is therefore bounded, and the fraction captured by a given instrument is bounded again. This is the photon budget. Everything else is a question of how to spend it. Sampling: pixels, voxels and Nyquist An image is not only blurred and noisy; it is also discretised into pixels, and in three dimensions into volume elements called voxels. The Nyquist–Shannon sampling theorem states that to capture all the information in a band-limited signal, it must be sampled at a spacing no larger than half the period of its highest frequency component. The microscope's PSF acts as a low-pass filter, removing spatial frequencies above a cutoff of 2NA/λ in the lateral plane. The Nyquist criterion therefore requires a pixel size of at most λ/(4NA) at the specimen. For the 63×, NA 1.4 objective in Table 1, that works out to about 93 nanometres. In practice many microscopists aim for pixels of around 80 to 100 nanometres at the specimen for high-NA imaging, and slightly finer sampling is often used when deconvolution is planned. Sampling too coarsely, called undersampling, throws away resolution the optics provided and can introduce aliasing artefacts, in which fine periodic structures appear as spurious coarser patterns. Sampling too finely, oversampling, wastes photons: the same number of photons spread over more pixels leaves each pixel noisier, and in a point-scanning system each extra pixel costs extra exposure time and extra excitation. This is the first appearance of a theme that will recur throughout. There is no free lunch in resolution. Every choice that captures finer detail does so by spreading a fixed supply of photons more thinly, and the noise per measurement rises accordingly. The four-way trade-off It is now possible to state the central trade-off precisely. A fluorescence experiment must balance four quantities: spatial resolution, temporal resolution (how fast images are acquired), signal-to-noise ratio, and specimen viability or, for fixed samples, the total number of photons the labels can yield before bleaching. These are coupled through the photon budget. Higher spatial resolution requires finer sampling, which requires more pixels, each of which needs enough photons to rise above noise. Faster acquisition means fewer photons per frame unless excitation intensity is raised, which accelerates bleaching and phototoxicity. Longer time-lapse series spread the photon budget across more frames. Deeper imaging in tissue loses photons to scattering and aberration, requiring more excitation to compensate. Every one of these trades against every other. Microscopists sometimes draw this as a pyramid or tetrahedron with one quantity at each vertex, and note that improving any one moves the experiment away from the others. The image is helpful, but it understates one point: the trade-offs are not symmetrical. Photobleaching and phototoxicity typically scale faster than linearly with excitation intensity in many regimes, because high intensities drive fluorophores into reactive excited states and multiphoton processes. Spreading the same total dose over a longer period at lower intensity is often markedly gentler than delivering it quickly at high intensity. This nonlinearity is one reason light-sheet microscopy, which illuminates each plane at modest intensity, is so much gentler than point-scanning confocal microscopy for the same total signal, a comparison developed in Chapter 4. Beyond the limit: three strategies If the diffraction limit is set by the physics of focusing, how can any technique beat it? The answer is that the Abbe limit applies strictly to linear imaging of a specimen whose fluorescent molecules all behave identically and simultaneously. Relax either assumption and the limit loosens. The first strategy is to engineer the illumination so that the effective PSF is smaller than the diffraction-limited one. Confocal microscopy does this modestly, by multiplying an illumination PSF by a detection PSF, as Chapter 2 explains. Structured illumination microscopy (SIM) does it by projecting fine patterns onto the specimen and computationally extracting the high-frequency information those patterns shift into the observable range; in its standard linear form, introduced by Mats Gustafsson in 2000, it doubles resolution in each dimension. These methods improve resolution by a bounded factor, typically up to two, and they remain within the regime of linear optics. The second strategy uses a nonlinear response of the fluorophore to confine emission. If a molecule's fluorescence can be switched off by light, and the off-switching saturates, then illuminating with a pattern that has a zero at its centre leaves only a sub-diffraction region in the "on" state. This is the principle of STED and its relatives, covered in Chapter 6. Because the confinement depends on how strongly the switching saturates, resolution in principle has no fixed lower bound. The third strategy separates molecules in time. If only a sparse subset of fluorophores emits in any given frame, each can be localized individually with precision far better than the PSF width, and a super-resolved image can be assembled from many thousands of frames. This is the basis of PALM, STORM, DNA-PAINT and related methods covered in Chapter 7. All three strategies share a feature that follows directly from the square-root law. Resolution below the diffraction limit is purchased with photons. SIM needs multiple raw images per reconstructed frame and high signal-to-noise ratio in each. STED needs intense depletion light that stresses fluorophores. Localization microscopy needs thousands of photons from each of many thousands of molecules. The diffraction limit was never a wall so much as a price point, and the techniques in this book offer different ways of paying above it. What the numbers do not show Before turning to the instruments, one caution. The resolution formulas in this chapter describe ideal optics, matched immersion media, and specimens that do not scatter or aberrate the light. Real specimens do all of these things. A cell mounted in aqueous medium beneath an oil objective introduces spherical aberration that grows with depth; ten microns into such a sample the axial PSF may already be noticeably elongated and dimmed. Tissue scatters visible light so strongly that ballistic, unscattered photons fall off exponentially with depth. Refractive index variations within the specimen itself, between cytoplasm, nuclei, lipid droplets and extracellular matrix, distort the wavefront in ways no fixed optic can correct. These imperfections mean that the practical resolution of an experiment is frequently set not by the objective's NA but by the sample. Chapter 8 returns to this problem and to the adaptive optical methods that address it. For now, the point is that a microscopist should treat the numbers in Table 1 as upper bounds, rarely reached in a biological specimen, and should measure the PSF of their own system under their own conditions, using sub-resolution fluorescent beads embedded in a medium that mimics the sample, whenever resolution matters to the conclusion. The instruments described in the following chapters are all responses to the same pair of constraints. What distinguishes them is where they choose to accept a cost, and what they buy with it. Chapter 2. Confocal Microscopy: Buying Clarity by Discarding Light A widefield fluorescence microscope illuminates the whole thickness of the specimen at once. Every fluorophore in the cone of illumination is excited, and every one of them contributes light to the camera image. Molecules in the focal plane contribute sharp images; molecules above and below it contribute broad, dim haze. In a thin cultured cell this haze is tolerable, and widefield imaging combined with deconvolution can do remarkably well. In a thick specimen, a tissue slice, an organoid, an embryo, the out-of-focus light can overwhelm the in-focus signal entirely, leaving an image in which structures are visible only as smudges on a bright fog. The confocal microscope solves this problem with a simple and rather wasteful idea: it refuses to record light that did not come from the focal point. Its success over four decades as the standard instrument for three-dimensional fluorescence imaging rests on that refusal. Its limitations, which have driven much of the development described in later chapters, rest on the same refusal. The pinhole and optical sectioning Marvin Minsky, then a junior fellow at Harvard, designed and built the first confocal microscope in the mid-1950s, motivated by a wish to trace neural connections in thick brain tissue. He filed a patent in 1957, granted in 1961, describing an instrument that illuminated the specimen with a focused point of light and detected the returning light through a pinhole placed in a plane conjugate to that point. The word "confocal" refers to this arrangement: the illumination focus and the detection pinhole are focused on the same point in the specimen. The physics of the pinhole is easy to state. Light emitted from the focal point is brought to a sharp focus at the pinhole and passes through it. Light emitted from a point above or below the focal plane comes to a focus in front of or behind the pinhole plane, so at the pinhole it is spread over a disc, and only a small fraction of it passes. The further a source lies from the focal plane, the larger that disc and the smaller the fraction transmitted. Out-of-focus light is not eliminated, but it is strongly suppressed. The result is optical sectioning: the ability to record an image of a thin slice within a thick specimen without physically cutting it. To build an image, the focused spot must be moved across the specimen point by point, and the signal from each point recorded as one pixel. Minsky moved the specimen stage mechanically. Practical instruments became widespread only in the 1980s, when laser sources, galvanometer mirrors for beam scanning, and computers for image storage came together; the confocal systems developed by Brad Amos, John White and colleagues at the Medical Research Council Laboratory of Molecular Biology in Cambridge, commercialised by Bio-Rad in 1987, were particularly influential. The pinhole size The pinhole diameter is usually expressed in Airy units, where one Airy unit (AU) is the diameter of the first dark ring of the Airy pattern projected onto the pinhole plane. A pinhole of 1 AU passes most of the in-focus light, about 84 per cent of the energy in an ideal Airy disc falls within the first dark ring, while providing good sectioning. Most instruments default to this value, and for good reason. Closing the pinhole below 1 AU improves sectioning further and, in principle, improves lateral resolution. The theoretical basis for the lateral gain is that the effective PSF of a confocal system is the product of the illumination PSF and the detection PSF. Multiplying two similar bell-shaped functions yields a narrower one; for Gaussian approximations the width shrinks by a factor of the square root of two. In the limit of an infinitely small pinhole, the confocal lateral resolution is therefore about 1.4 times better than widefield. In practice this improvement is almost never realised, because a vanishingly small pinhole rejects almost all the light. At 0.2 to 0.3 AU, where much of the resolution gain appears, the pinhole passes only a small fraction of the emitted photons. The signal-to-noise ratio collapses, and the square-root law of Chapter 1 ensures that recovering it requires a large increase in exposure, which accelerates bleaching. A confocal microscope operated at 1 AU achieves lateral resolution only modestly better than widefield, perhaps five to ten per cent under realistic conditions. Its genuine advantage is sectioning, not lateral resolution. This is worth stating plainly, because the belief that confocal microscopy substantially improves lateral resolution is widespread and leads researchers to close pinholes in pursuit of a gain that costs far more than it returns. Opening the pinhole beyond 1 AU, conversely, increases signal at the expense of sectioning. For dim live samples where sectioning is less important than survival, a pinhole of 1.5 or 2 AU is often the better choice. The pinhole, in other words, is a direct control over the exchange rate between photons and optical sectioning. Point scanning and its costs A point-scanning confocal microscope builds each image one pixel at a time. At any instant the entire excitation power is concentrated into a diffraction-limited spot, and the time spent on each pixel, the pixel dwell time, is short: typically between a fraction of a microsecond and a few microseconds. A 1024 by 1024 image at one microsecond per pixel takes about a second to acquire, once line flyback is included. This architecture has three consequences that shape its use. The first is speed. Conventional galvanometer mirrors scan a line in roughly a millisecond, giving frame rates of a few per second at moderate image sizes. Resonant scanners, which oscillate at a fixed frequency of around 8 or 12 kilohertz, achieve video rates or faster, but at the cost of very short pixel dwell times, of the order of tens of nanoseconds, so that each pixel receives only a handful of photons per sweep. Averaging multiple sweeps recovers signal but surrenders the speed. The second is peak intensity. To collect enough photons in a microsecond, the focal intensity must be very high, often in the range of kilowatts to hundreds of kilowatts per square centimetre. At these intensities fluorophores approach saturation: a molecule spends a significant fraction of each dwell time in its excited or triplet state and cannot absorb further photons. Raising laser power then yields diminishing signal while continuing to drive photochemistry. Saturation is one of the reasons point-scanning confocal microscopy is harsher on live samples than its average light dose would suggest. The same number of photons delivered at lower intensity over a longer time is, for many fluorophores, gentler. This is the nonlinearity mentioned in Chapter 1, and it favours parallelised illumination wherever the application allows. The third consequence is that the whole specimen volume in the illumination cone is excited even though only the focal point is recorded. Every time the focal plane is imaged, fluorophores above and below it are exposed to excitation light and bleach, although their emission is rejected by the pinhole. In a z-stack of fifty planes, each plane receives excitation during the imaging of every other plane. For thick live specimens this out-of-focus excitation is a major and invisible component of the photon cost, and it is the central motivation for both two-photon excitation and light-sheet illumination. Spinning disks: parallelising the pinhole The spinning-disk confocal microscope addresses speed and peak intensity by using many pinholes at once. The idea goes back to the Nipkow disk, a spiral array of holes patented by Paul Nipkow in 1884 for early television. Mojmír Petráň, Milan Hadravský and colleagues in Czechoslovakia built a tandem-scanning reflected-light microscope based on a Nipkow disk in the late 1960s. For fluorescence, the decisive development was the microlens-enhanced design commercialised by Yokogawa in the 1990s, in which a second disk carrying thousands of microlenses focuses the excitation light through the corresponding pinholes of the first. This raised the fraction of excitation light passing the disk from a few per cent to a much larger value and made spinning-disk systems practical for fluorescence. As the disk rotates, the array of pinholes sweeps across the field, and a camera integrates the emission over the exposure. Each point in the specimen is illuminated many times per frame, briefly, at intensities far lower than a single point-scanning spot. Frame rates of tens to hundreds per second become possible, limited mainly by the camera and the signal, and the lower peak intensity makes photobleaching per detected photon generally lower than in point scanning for many live-cell applications. The costs are two. The pinhole size is fixed by the disk and matched to a particular objective magnification, so the photon-sectioning exchange rate cannot be adjusted freely. And in thick specimens, out-of-focus light emitted near one pinhole can pass through a neighbouring pinhole, a phenomenon called pinhole crosstalk, which degrades sectioning compared with a point scanner. Disks with wider pinhole spacing reduce crosstalk at the cost of light throughput. As a result, spinning-disk confocals are the instrument of choice for fast imaging of cells and thin tissues up to a few tens of microns, while point scanners remain preferable for thick, scattering specimens where sectioning quality matters most. Detectors: where modern confocal performance lives For much of the history of the confocal microscope, the detector was the weakest link. Conventional photomultiplier tubes with multialkali photocathodes have quantum efficiencies of around 20 to 30 per cent in the green, and their gain process adds multiplicative noise. The introduction of GaAsP photocathodes raised quantum efficiency in the visible to around 40 to 45 per cent, roughly doubling the photons recorded for the same excitation. Hybrid detectors, which combine a GaAsP photocathode with an avalanche diode in place of the dynode chain, add a further advantage: their gain is much less noisy, so that individual photon arrivals produce pulses of nearly uniform height and can be counted digitally. Photon counting changes the character of confocal data. Instead of an analogue voltage whose relationship to photon number depends on gain settings, offset and detector noise, each pixel contains an integer count of detected photons. The Poisson statistics of Chapter 1 then apply directly, which is a considerable advantage for quantitative work and for deconvolution. Counting has a limit at high rates, where closely spaced photons are missed during the detector's dead time, but for most biological samples imaged at sensible intensities this limit is not reached. Pulsed lasers combined with time-resolved photon counting also make it possible to record when each photon arrives relative to the excitation pulse. This is the basis of fluorescence lifetime imaging, which distinguishes fluorophores or environments by their decay times rather than their colours, and of time-gated detection, which discards photons arriving in the first nanosecond or so after excitation. Gating is useful for rejecting reflected light and short-lived autofluorescence, and, as Chapter 6 explains, it is an important component of modern STED. Image scanning microscopy: keeping the light the pinhole threw away The pinhole trade-off described earlier seems unavoidable: small pinholes give resolution but lose light, large pinholes keep light but lose resolution. A more recent insight shows that the trade-off is partly an artefact of using a single detector. Colin Sheppard pointed out in 1988 that if the pinhole is replaced by an array of small detectors, each element sees the specimen through a slightly displaced tiny pinhole. Each element therefore records an image with the improved resolution of a closed-pinhole confocal, but shifted by an amount that depends on its position. The signal that falls outside the central element is not noise to be rejected but a set of offset high-resolution images. Reassign each photon to the position midway between the illumination spot and the point that detector element observes, and sum the results, and the resolution benefit of a small pinhole is obtained with the light collection of a large one. Claus Müller and Jörg Enderlein demonstrated this experimentally in 2010 under the name image scanning microscopy. Commercial implementations followed quickly. Zeiss's Airyscan detector, introduced in 2014, uses a hexagonal array of 32 GaAsP elements behind a fibre bundle in the pinhole plane; with reassignment and subsequent deconvolution it improves lateral resolution by up to about 1.7-fold relative to widefield, to roughly 120 to 140 nanometres, while collecting substantially more light than a closed pinhole. Other approaches, including re-scan confocal microscopy and optical photon reassignment in instant SIM (introduced by Andrew York and colleagues in 2013), perform the reassignment optically rather than computationally, enabling faster acquisition. These methods sit on the boundary between confocal microscopy and structured illumination, and they illustrate a lesson that recurs throughout advanced microscopy. Photons that a traditional design discards frequently carry information. A cleverer detector, or a cleverer use of the data, can recover it. The reassignment approaches are now standard on many high-end confocals and represent, for many laboratories, the most accessible form of modest super-resolution: roughly a 1.5- to 1.7-fold improvement with conventional dyes and sample preparation, and without the extreme intensities of STED or the long acquisitions of localization microscopy. Colour: spectral detection and crosstalk Most confocal experiments image more than one label, and the way a system separates colours is another place where photons are either spent well or wasted. Traditional instruments used fixed dichroic mirrors and bandpass filters in front of each detector. Modern point scanners more often disperse the emission with a prism or grating and select detection windows with adjustable slits or mirrors, so that each channel's window can be tuned to the emission spectrum of a particular dye. The flexibility is valuable, but the principle behind it is the same as with filters: each window is a compromise between collecting as much of one label's emission as possible and excluding the emission of its neighbours. Fluorophore emission spectra are broad and asymmetric, with long red tails. A green dye such as fluorescein or Alexa Fluor 488 still emits a measurable fraction of its light in the orange range where a red dye is detected, a phenomenon called bleed-through or crosstalk. Excitation crosstalk compounds the problem: the laser line intended for one dye often excites another weakly. In a colocalisation experiment, where the question is whether two proteins occupy the same structures, uncorrected bleed-through manufactures colocalisation that is not there. The standard defences are sequential acquisition, in which each laser line is paired with its own detection window and the channels are recorded one after another, and single-label controls, which measure how much each dye contributes to each channel. When spectra overlap too much for windows to separate them, linear unmixing offers a computational alternative. The emission spectrum of each dye is measured in reference samples, the specimen is recorded in many narrow spectral channels, and each pixel's spectrum is decomposed into a weighted sum of the references. Unmixing can separate dyes whose peaks lie only a few nanometres apart and can remove broad autofluorescence as if it were an additional label. It is not free, however. Dividing the emission into many narrow channels divides the photons with it, and the decomposition amplifies noise when the reference spectra are similar. The same trade appears again: separation that the optics cannot provide is bought computationally, and the price is paid in signal-to-noise ratio. When confocal is the wrong tool The confocal microscope's great virtue is versatility. It handles fixed and live samples, thin and moderately thick specimens, multiple colours with flexible detection windows, and a range of quantitative measurements such as fluorescence recovery after photobleaching and colocalisation. For fixed samples up to perhaps fifty to a hundred microns thick, it remains in most laboratories the default. Its limits are nonetheless clear once the photon budget is considered. In highly scattering tissue, beyond roughly one or two scattering lengths, the returning light that has been scattered no longer comes to a focus at the pinhole and is rejected along with genuine out-of-focus light; the signal falls exponentially with depth, and the excitation spot itself degrades. This is the regime in which two-photon excitation, discussed next, holds the advantage. In long time-lapse imaging of whole embryos or organoids, the repeated out-of-focus excitation of a point-scanning system consumes the budget of the entire volume to image one plane at a time; light-sheet microscopy, discussed in Chapter 4, avoids this cost. And where resolution below roughly 120 nanometres is required, no confocal design will suffice, and the methods of Chapters 6 and 7 become necessary. It is worth closing with the practical settings that most often waste photons on a confocal system, because they are the ones a new user controls. Pinholes closed below 1 AU in pursuit of resolution that will not appear. Pixel sizes far smaller than Nyquist sampling requires, so that each pixel receives too few photons to be useful. Laser power raised to compensate for a detector gain that could have been increased instead, or for a detector that would have been better operated in photon-counting mode. Line averaging used to smooth noise when a slower scan at lower power would deliver the same photons more gently. None of these errors is exotic, and all of them spend the specimen's photon budget without buying anything the experiment needs. Chapter 3. Two-Photon Excitation: Spending Laser Power to Buy Depth Biological tissue is turbid. Light entering the brain, a lymph node or a tumour is deflected by countless small refractive index boundaries, between membranes, organelles, myelin and extracellular fluid, so that after travelling a short distance most photons are no longer heading in the direction they started. A confocal microscope depends on unscattered, or ballistic, light in both directions: excitation must reach the focus without deviation, and emission must return along the same path to pass the pinhole. As depth increases the ballistic fraction falls exponentially, and beyond roughly a hundred microns in most tissues confocal imaging becomes impractical. Two-photon excitation microscopy changed this. It does not stop scattering; it makes scattering matter less. By confining excitation to the focal volume through a nonlinear optical process, it removes the need for a pinhole, so that every emitted photon, scattered or not, can be counted as signal. And by using near-infrared light, which scatters less than visible light, it delivers excitation deeper. The combination made it possible to image living neurons several hundred microns beneath the surface of the brain, and it has become the standard method for functional imaging of neural activity in vivo. The physics of simultaneous absorption In ordinary one-photon fluorescence, a molecule absorbs a single photon whose energy matches the gap between its ground and excited states. In two-photon excitation, the molecule absorbs two photons at essentially the same moment, each carrying about half the required energy. The combined energy lifts the molecule to the same excited state, from which it fluoresces normally. A dye excited at 488 nanometres in one-photon mode can therefore be excited by two photons at roughly twice that wavelength, in the near infrared around 900 to 1000 nanometres. Maria Göppert-Mayer predicted the process theoretically in her 1931 doctoral thesis, and the unit of two-photon absorption cross-section, the GM, is named after her. Observing it required light intensities that became available only with lasers. Its application to microscopy was demonstrated by Winfried Denk, James Strickler and Watt Webb at Cornell University in a paper in Science in 1990. The property that makes two-photon excitation useful is its intensity dependence. Because two photons must arrive within a very short interval, of the order of a femtosecond, the probability of absorption is proportional to the square of the instantaneous light intensity. When a laser beam is focused by a high-NA objective, intensity is enormous at the focus and falls off rapidly away from it: in the paraxial picture, the beam's cross-sectional area grows with the square of distance from focus, so intensity falls with the inverse square. With a quadratic dependence, excitation per plane falls off with the inverse square of the distance, and integrated over each plane the total excitation away from focus becomes small. Almost all the fluorescence is generated within a small ellipsoid at the focal point, typically under a micron wide and a few microns long for a high-NA objective. This intrinsic confinement has two consequences, and together they are the whole case for the technique. First, there is no out-of-focus excitation, and therefore no out-of-focus bleaching or phototoxicity. The planes above and below the focus are traversed by the infrared beam but are not excited, which is a direct answer to the hidden cost of confocal imaging described in Chapter 2. Second, every emitted photon must have come from the focus, so no pinhole is required to reject out-of-focus light. The detector can be placed to collect as much emission as possible, including emission that has been scattered on its way back out of the tissue. Pulsed lasers and the peak-power trick The quadratic dependence carries a practical problem. At the average powers that tissue can tolerate, tens to perhaps a couple of hundred milliwatts, a continuous laser would produce negligible two-photon absorption. The solution is to concentrate the light in time. Mode-locked titanium-sapphire lasers, which became commercially available in the early 1990s, emit pulses around 100 femtoseconds long at a repetition rate of about 80 megahertz. The laser is therefore on for only about one hundred thousandth of the time, and the peak intensity during each pulse is about a hundred thousand times the average. Because two-photon absorption depends on the square of intensity, the time-averaged excitation rises by roughly the same factor compared with a continuous beam of the same average power. This duty-cycle arithmetic explains many practical details. Pulse duration matters: the glass in an objective and other optics stretches femtosecond pulses through group-velocity dispersion, and a pulse broadened from 100 to 200 femtoseconds yields roughly half the excitation. Commercial systems therefore often include prism or grating compressors to pre-compensate. Repetition rate matters too: reducing it while keeping average power constant raises the energy per pulse and the signal per pulse, which is the basis of low-repetition-rate sources used for deep imaging. And tunability matters, because different fluorophores have different optimal two-photon excitation wavelengths, and these do not always sit at exactly twice their one-photon peaks. Two-photon absorption spectra are often broad and blue-shifted relative to the doubled one-photon spectrum, which allows several fluorophores to be excited simultaneously by a single wavelength, a convenience for multicolour work. The cost of this approach is that the lasers are large, expensive and complex compared with the continuous diode lasers used in confocal microscopy. Fixed-wavelength fibre lasers at around 920 and 1040 nanometres have reduced both cost and complexity, and these two wavelengths conveniently match green calcium indicators and red fluorescent proteins respectively. Why near-infrared light goes deeper Scattering in tissue decreases as wavelength increases. The characteristic distance over which ballistic light is attenuated by a factor of e, the scattering mean free path, is roughly 50 to 100 microns for visible light in mammalian cortex and on the order of 150 to 200 microns for near-infrared light around 900 nanometres, although exact values vary with tissue, age and preparation. Two-photon excitation therefore delivers its excitation beam into tissue with less loss than a visible beam. There is a further, subtler advantage. Scattered excitation photons in a two-photon microscope are spread out in space and time and have very low intensity, so they produce essentially no two-photon absorption. The focus is formed only by the ballistic portion of the beam, and the excitation remains confined even though much of the light has been scattered. In a one-photon system, scattered excitation light still excites fluorophores wherever it goes, generating background. On the detection side, the emitted fluorescence is visible light and scatters heavily on its way out. In a confocal microscope, this scattered emission would miss the pinhole. In a two-photon microscope, it need not be discarded. The most effective designs use non-descanned detectors: large-area photomultipliers placed close to the back of the objective, collecting emitted light directly without sending it back through the scanning mirrors. Because the origin of every photon is already known from the position of the scanned focus, the path by which the photon arrives is irrelevant. This is perhaps the most elegant feature of two-photon microscopy: it converts the scattered emission that defeats confocal imaging into usable signal. The depth limit How deep can two-photon imaging go? Excitation at the focus falls with depth as the ballistic fraction of the beam decreases, and this can be compensated by raising laser power, up to limits set by tissue heating and damage. But as power rises, a second problem emerges. Near the surface of the tissue, where the beam is broad but still intense, a small amount of two-photon excitation occurs over a large volume. At shallow depths this out-of-focus background is negligible compared with the focal signal. At great depths, where the focal signal has been attenuated, it is not. Patrick Theer, Mazahir Hasan and Winfried Denk showed in 2003 that this surface-generated background sets a fundamental depth limit, and using a regenerative amplifier with high pulse energy they imaged to around one millimetre in the mouse neocortex. For typical labelling and standard lasers, the practical limit in mouse cortex is usually quoted as roughly 500 to 800 microns, reaching the upper cortical layers but not the deeper layers or the hippocampus beneath without removing overlying tissue. Three-photon excitation: extending the reach The logic that makes two photons better than one extends to three. In three-photon excitation, a fluorophore absorbs three photons simultaneously, each carrying roughly a third of the required energy. Absorption now depends on the cube of intensity, which confines excitation still more tightly and, crucially, suppresses the out-of-focus background from near the surface that limits two-photon depth. Three-photon excitation of a green fluorophore requires light around 1,300 nanometres; a red one, around 1,700 nanometres. Both wavelengths sit in windows where tissue scattering is lower than at 900 nanometres, and where absorption by water, which rises sharply at longer infrared wavelengths, has local minima. Nicholas Horton, Chris Xu and colleagues at Cornell demonstrated three-photon imaging of subcortical structures in the intact mouse brain in 2013, including red-labelled neurons in the hippocampus, beyond the reach of two-photon methods. Subsequent work has extended three-photon functional imaging of neural activity to depths beyond one millimetre in mouse brain. The costs are considerable. Three-photon cross-sections are extremely small, so pulse energies must be much higher, and lasers typically operate at low repetition rates of around one to a few megahertz to deliver adequate energy per pulse without exceeding average-power limits. Fewer pulses per second means fewer excitation events per pixel, so frame rates and fields of view are more limited than in two-photon imaging. Three-photon imaging thus illustrates the budget framework neatly: a still more nonlinear process buys still greater depth, and pays for it in speed, signal and equipment cost. The costs of speed and power The scanning bottleneck Because two-photon excitation happens at a single focal point, a two-photon microscope inherits the speed limitations of point scanning. A resonant scanner running at about 8 kilohertz can produce 512-line frames at around 30 per second, which is adequate for calcium signals in a single plane, where indicator kinetics are themselves slow on the scale of tens to hundreds of milliseconds. But neural circuits are three-dimensional, and sampling a volume by stepping the focus through successive planes divides the frame rate by the number of planes. At each pixel the dwell time is already only tens of nanoseconds, a few pulses of the laser, and it cannot be shortened much further without the fluorescence lifetime itself, a few nanoseconds, becoming the limit. Several strategies spend the photon budget differently to escape this bottleneck. One is to elongate the focus deliberately, for example by shaping the excitation into a Bessel-like beam tens of microns long along the axis, so that a single two-dimensional scan projects an entire volume onto one image; Rongwen Lu, Na Ji and colleagues used this approach in 2017 to record synaptic-scale activity in volumes at video rate, accepting the loss of axial information in exchange for speed in sparsely labelled tissue. Another is multiplexing: splitting each laser pulse into several beamlets focused at different depths and delayed in time by a few nanoseconds, so that the fluorescence from each can be separated by its arrival time at a single detector. Light-beads microscopy, reported by Jeffrey Demas, Alipasha Vaziri and colleagues in 2021, pushes this idea to dozens of axially separated foci, enabling volumetric recording across large regions of mouse cortex. A third approach abandons the idea of scanning every pixel at all. If the positions of the neurons of interest are known, the laser can be directed only to them, either by random-access scanning with acousto-optic deflectors, which can jump between arbitrary points in microseconds, or by holographic patterning with a spatial light modulator, which can illuminate many cell bodies at once. These targeted methods waste no photons on the neuropil between cells, and they are closely related to the holographic techniques used to stimulate chosen neurons optogenetically while recording from others. Each of these strategies makes a specific bet about what information is dispensable, whether axial detail, background pixels or the space between cells, and redirects the saved excitation to speed or volume. None is universally better than conventional raster scanning. They are better when the bet is right for the question. Heating and photodamage The photon budget in two-photon microscopy is enforced by two distinct damage mechanisms, and it is important to distinguish them. The first is heating. Near-infrared light is absorbed, mostly by water and haemoglobin, and the absorbed energy raises tissue temperature. Because this absorption is linear, it depends on average power rather than on peak intensity. Kaspar Podgorski and Gayathri Ranganathan measured and modelled brain heating during multiphoton imaging in a 2016 study in the Journal of Neurophysiology, and showed that the laser powers commonly used for deep imaging in mouse cortex, of the order of a couple of hundred milliwatts at the brain surface, can raise local temperature measurably, and that sustained illumination at the upper end of commonly used powers produced physiological effects and signs of tissue damage. Their results translated a vague sense that "too much power is bad" into a quantitative guide, and power limits in the range of roughly 150 to 250 milliwatts are now commonly cited for sustained imaging with standard objectives and fields of view, although the safe figure depends on duty cycle, field size and wavelength. The second mechanism is nonlinear photodamage at the focus. Because excitation depends on intensity squared, and higher-order processes on still higher powers, raising the peak intensity drives multiphoton ionisation and the formation of reactive species within the focal volume. Experiments on cells and tissue have found that this form of damage rises more steeply with power than the fluorescence signal does; Hopt and Neher reported in 2001 that photodamage in their preparation scaled with roughly the 2.5th power of laser intensity. The practical consequence is that it is generally better to increase signal by collecting more efficiently, using better detectors or brighter indicators, than by increasing peak power, since the damage rises faster than the reward. These two mechanisms pull in opposite directions when choosing laser parameters. Lowering repetition rate at fixed average power raises peak intensity, increasing signal but also nonlinear damage. Raising average power at fixed repetition rate increases both signal and heating. A careful experimenter treats both as limits and optimises within the region they bound, rather than simply turning the laser up until the image looks good. What two-photon imaging made possible The clearest case for two-photon microscopy is in neuroscience. Combined with genetically encoded calcium indicators, notably the GCaMP family, whose GCaMP6 variants were described by Tsai-Wen Chen and colleagues at the Janelia Research Campus in 2013, two-photon imaging allows the activity of hundreds or thousands of individual neurons to be recorded simultaneously in an awake animal performing a task. Large-field instruments such as the two-photon mesoscope described by Nicholas Sofroniew, Karel Svoboda and colleagues in 2016 extended the field of view to several millimetres, capturing activity across multiple cortical areas. Beyond neuroscience, two-photon imaging has been used to follow immune cells migrating through lymph nodes, to study tumour cell invasion in living animals, and to image kidney function and skin. It also enables label-free contrast: second-harmonic generation from non-centrosymmetric structures such as collagen fibrils and microtubule bundles can be recorded with the same laser and detectors, giving a view of tissue architecture without any stain. Two-photon microscopy has its weaknesses. Its resolution is somewhat worse than confocal at the same NA, because the effective excitation wavelength is roughly double; the quadratic dependence narrows the effective PSF but not enough to fully compensate. It is slower than camera-based methods, because it remains a point-scanning technique, and the scanning must be fast enough to capture neural dynamics, which constrains the pixel dwell times and hence the photons per pixel. Its equipment is costly. And for thin specimens it offers little advantage over confocal imaging, since the scattering it overcomes is not present. But within its domain, which is deep imaging in living, scattering tissue, it has had no real competitor for three decades. It is the clearest example in microscopy of the budget framework: a technique that spends expensive, carefully shaped infrared power, accepts a modest loss of resolution, and buys in return the one thing most other techniques cannot supply, which is sight beneath the surface. Hashtags: #AdvancedMicroscopy #FluorescenceMicroscopy #ConfocalMicroscopy #TwoPhotonMicroscopy #SuperResolutionMicroscopy #PhotonBudget #DiffractionLimit #PointSpreadFunction #OpticalSectioning #Photobleaching #Phototoxicity #LightSheetMicroscopy #STEDMicroscopy #PALM #STORM #DNAPAINT #MINFLUX #StructuredIlluminationMicroscopy #AdaptiveOptics #Deconvolution #ImageScanningMicroscopy #FluorescenceLifetimeImaging #ThreePhotonMicroscopy #QuantitativeImaging #FutureOfAdvancedMicroscopy
- Translational Science Fundamentals (Moving Discoveries from Bench to Bedside)
Download the Book (PDF): Introduction In 1983, researchers publishing in the most prestigious basic-science journals in the world were confident about the clinical promise of their work. Two decades later, a team led by Despina Contopoulos-Ioannidis and John Ioannidis went back and checked. They identified 101 articles from leading journals, published between 1979 and 1983, whose authors had explicitly claimed that a discovery held real promise for prevention or treatment. By 2002, only 27 of those discoveries had been tested in a randomized trial. Five had led to a licensed product. One was in wide clinical use. The finding, published in the American Journal of Medicine in 2003, is not an indictment of basic science. Most of those papers were good science. It is a measurement of how much happens, or fails to happen, between a discovery and a patient. A second number circulates in almost every lecture on this subject. In 2000, Andrew Balas and Suzanne Boren estimated that it takes about seventeen years for research evidence to reach routine clinical practice, and that only about fourteen percent of original research ever does. The figure has been criticised, refined and occasionally mocked, most usefully by Zoë Morris, Steven Wooding and Jonathan Grant, whose 2011 review in the Journal of the Royal Society of Medicine was titled "The answer is 17 years, what is the question." Their point was that the time lag depends entirely on where one starts the clock and where one stops it, and that different studies measure different segments of a long and branching path. That point is the beginning of this book. What "translation" means Translational science is the study of how knowledge moves from one kind of work to another: from a laboratory observation to a first human dose, from a trial result to a clinical guideline, from a guideline to what actually happens in a clinic on a Tuesday afternoon, and from an effective intervention to a change in the health of a whole population. It is distinct from translational research, which is the work of moving any particular discovery along that path. The science asks why the path is so slow, so lossy and so uneven, and what can be done about it. Over the past two decades, the field has converged on a vocabulary of numbered stages. T0 is discovery: the identification of a mechanism, a target, a biomarker or a candidate intervention. T1 is the move into humans: first-in-human studies and the early trials that establish safety, dose and proof of mechanism. T2 establishes efficacy in patients and turns that evidence into guidance. T3 is the move from guidance into practice: implementation, dissemination and the study of why proven interventions are or are not used. T4 is the move from practice to population: the outcomes that follow at scale, and the policies that shape whether an intervention reaches everyone who could benefit. Different institutions draw these lines in slightly different places, and the history of how the lines came to be drawn is itself instructive. What matters more than the precise boundaries is the recognition that each stage asks a different question, demands a different kind of evidence, is carried out by different people in different institutions, and is funded and rewarded by different means. The argument of this book The central claim of what follows is simple to state and has consequences that take the rest of the book to work through. Discoveries do not stall in the pipeline so much as at its joints. The characteristic failure of translation is not that the science at any one stage is bad, though sometimes it is. It is that each stage is optimised for its own question and hands the next stage something it was never designed to use. A preclinical study optimised to publish a striking mechanism produces an effect size that no human trial will reproduce. A Phase I study optimised to find the highest tolerable dose hands forward a dose that may be wrong for the patients who will take it for years. An efficacy trial optimised for internal validity, with carefully selected participants and expert investigators, hands forward a result that community clinics cannot reproduce with their own patients and staff. An implementation programme optimised for uptake in one setting hands forward a model that no payer will fund and no legislature will mandate. At every joint, a question that the next stage needed answered was not asked, because asking it was nobody's job. The remedy, which the book develops stage by stage, is to design each stage with the next one's question in view. That means preclinical studies built to predict human outcomes rather than merely to demonstrate mechanism; first-in-human studies that collect the pharmacology later trials will need; efficacy trials that test interventions in forms that can actually be delivered; implementation studies that measure what policymakers will ask about; and policies designed so that their effects can be evaluated. It also means recognising that the pipeline is not a line but a loop. Observations at the bedside and in populations constantly send questions back towards the bench, and some of the most productive translation has run in that direction. How the book is organised Chapter 1 sets out the T0 to T4 framework, its origins and its limits, and argues for reading it as a map of handoffs rather than a sequence of stages. Chapter 2 examines the first and most notorious joint, between preclinical discovery and the decision to test in humans, where problems of reproducibility and model validity cause most candidate interventions to fail before or shortly after they reach a person. Chapter 3 is devoted to the first-in-human study itself: how starting doses are chosen, what went wrong in the most instructive Phase I disasters, and how dose-finding designs have evolved. Chapter 4 follows candidates through the long and costly work of establishing efficacy, where attrition is highest and where the choice of endpoints and trial designs determines whether a result will mean anything outside the trial. The second half of the book turns from products to people and systems. Chapter 5 addresses the gap between evidence and practice, and the implementation science that has grown up to study it. Chapter 6 follows interventions into communities, where the tension between fidelity and adaptation plays out and where sustainability is decided. Chapter 7 examines T4, the translation of evidence into public health policy and the population effects that follow. Chapter 8 considers the pipeline as a system: the flows that run backwards from practice to discovery, the de-implementation of practices that should never have been adopted, and the infrastructure that institutions have built to connect the stages. The conclusion draws out what the whole analysis implies for the people who fund, conduct and use translational research. A note on scope Translational science spans drugs, biologics, devices, diagnostics, behavioural interventions and policies. This book draws examples from all of them but leans on two families of cases: pharmaceutical development, where the early stages are best documented, and public health prevention, where the late stages are. It is written from a largely American and British vantage point, because those are the systems whose translational infrastructure is most thoroughly studied, but the problems it describes are general. Wherever a figure or a study is cited, it is a real one, named so that the reader can find it. Where the evidence is contested, the book says so. The reader need not be a scientist. Anyone who has wondered why a promising headline about a new treatment so rarely becomes a treatment, or why a proven intervention sits unused for a decade, is asking the questions this field exists to answer. The answers turn out to be less about any single failure of intelligence or effort than about the structure of the enterprise itself, and structures, unlike laws of nature, can be redesigned. Chapter 1. The Pipeline and Its Joints The word "translational" entered the vocabulary of biomedical policy in the early 2000s, and it arrived as a diagnosis. In 2003, Nancy Sung and colleagues, writing for the Clinical Research Roundtable convened by the US Institute of Medicine, published an analysis in JAMA of the central challenges facing the national clinical research enterprise. They identified two "translational blocks." The first lay between basic biomedical research and its application in human studies. The second lay between the results of clinical studies and their adoption in everyday practice and decision-making. The same year, Elias Zerhouni, then director of the US National Institutes of Health, launched the NIH Roadmap for Medical Research, which named "re-engineering the clinical research enterprise" as one of its three themes and set in motion the institutional changes that would later produce the Clinical and Translational Science Awards. The two-block model had the virtue of simplicity, and it captured something real. American biomedical research budgets had doubled between 1998 and 2003, and there was a growing unease that the flood of basic discovery was not producing a corresponding flow of new therapies or improvements in health. But the model also concealed a great deal. It grouped together, under "the second block," activities as different as running a randomised trial, writing a clinical guideline, persuading a clinician to follow it, and persuading a government to pay for it. As researchers from different disciplines began to look at the path in detail, the number of blocks began to multiply. From two blocks to five stages The first elaboration came from primary care. In 2007, John Westfall, James Mold and Lyle Fagnan published a short and influential piece in JAMA titled "Practice-based research—'Blue Highways' on the NIH Roadmap." Their argument was that the move from a successful trial to routine practice is not a single step. Trials are conducted largely in academic medical centres, with selected patients and expert staff. Most patients, however, are seen in community practices, and there is a distinct body of research, conducted in those practices, that asks whether and how trial findings work there. Westfall and colleagues labelled this T3 and called for practice-based research networks to carry it out. In 2008, Steven Woolf wrote in JAMA on "The meaning of translational research and why it matters." Woolf observed that the term was being used for two quite different enterprises. For laboratory scientists and the pharmaceutical industry, translation meant turning discoveries into drugs and devices, work that depended on molecular biology, animal models and early-phase trials. For health services researchers, translation meant ensuring that proven interventions actually reached patients, work that depended on epidemiology, behavioural science, organisational research and policy analysis. The two groups competed for the same money under the same label while needing entirely different expertise, infrastructure and timelines. Woolf's warning was that the first enterprise, being closer to the traditional interests of academic medicine, would tend to absorb the funds intended for both. The fullest version of the framework came from genomic medicine and epidemiology. In 2007, Muin Khoury and colleagues at the US Centers for Disease Control and Prevention, writing in Genetics in Medicine, described a continuum of translation research for genomic discoveries running from T1 through T4. In 2010, Khoury, Marta Gwinn and John Ioannidis extended it in the American Journal of Epidemiology in a paper titled "The emergence of translational epidemiology: from scientific discovery to population health impact." They added a T0 stage for the discovery research that precedes all application, and argued that epidemiology had roles to play at every stage, not just the last. The resulting five-stage scheme is now the most common reference point. As Table 1 sets out, each stage answers a distinct question, uses characteristic kinds of study, and ends with a handoff that the next stage depends on. Table 1. The five translational stages and their handoffs. Stage Central question Typical studies What it hands forward T0 Is there a mechanism or target worth pursuing? Laboratory, animal, genomic and observational discovery research A candidate intervention and a rationale T1 Is it safe and active in humans, and at what dose? First-in-human, Phase I, early Phase II, proof of mechanism A dose, a safety profile, early signals T2 Does it work in patients, and should it be recommended? Phase II and III trials, systematic reviews, guidelines Evidence of efficacy and a recommendation T3 Is it adopted and delivered well in real practice? Implementation, dissemination and effectiveness research A deliverable programme and its uptake T4 Does it improve population health, and for whom? Outcomes research, policy evaluation, surveillance Population impact and policy Other institutions draw the lines differently. The US National Center for Advancing Translational Sciences, established in December 2011, describes a spectrum from basic research through pre-clinical, clinical, clinical implementation and public health research, and deliberately depicts it without arrows so as to emphasise that each stage can feed any other. Some schemes fold Phase II and Phase III into T1; others place guideline development in T3. These are not trivial differences, because the labels determine which funders and which review panels consider an application. But they share the essential insight: the journey from discovery to population health has several distinct segments, and a problem that is solved in one segment can remain entirely unsolved in the next. Why the joints matter more than the stages It is tempting to read the T-stages as a production line, in which each station does its work and passes the product along. The metaphor misleads in two ways, and both matter for everything that follows. First, the stations are not run by the same organisation. On a factory line, a single management can see that the paint shop is producing parts that the assembly shop cannot use, and fix it. In translational science, T0 is conducted largely by academic laboratories funded by research councils and rewarded by publications in high-impact journals. T1 and T2 for drugs are conducted largely by companies, regulated by agencies such as the US Food and Drug Administration and the European Medicines Agency, and rewarded by market approval. T3 is conducted by health systems, professional societies and a relatively small community of implementation researchers, and is rewarded, if at all, by quality metrics and payment incentives. T4 is conducted by public health agencies and legislatures and is rewarded by political and electoral considerations. No single actor is responsible for the whole line, and none is rewarded for the quality of what it hands forward. Second, each stage defines success in its own terms, and those terms do not always serve the next stage. A preclinical paper succeeds if it is published and cited. A Phase I study succeeds if it identifies a dose that can be carried forward without unacceptable toxicity. An efficacy trial succeeds if it achieves statistical significance on its primary endpoint. An implementation programme succeeds if the intervention is adopted. A policy succeeds if it is enacted. Each of these definitions is reasonable in isolation. Each can be satisfied while failing the stage that follows. A published mechanism can be unreproducible. A tolerable dose can be far higher than necessary. A statistically significant result can depend on a population that looks nothing like the patients who will use the treatment. An adopted programme can be delivered so poorly that it has no effect. An enacted policy can go unevaluated. The consequence is that the most useful way to read the T-framework is as a map of handoffs, and to ask of each joint three questions. What does the receiving stage need to know? Who, at the sending stage, is responsible for finding it out? And what would it cost them to do so? When the answers are "a great deal," "nobody," and "more than they are rewarded for," the joint will leak. Much of the history of translational science since 2003 can be read as a series of attempts to change those answers, by building infrastructure that spans stages, by changing reporting and design standards, and by creating incentives for the sending stage to care about what happens downstream. The seventeen-year problem, measured properly The best-known summary statistic in the field, the seventeen-year lag estimated by Balas and Boren in 2000, was derived by adding together estimates of the time taken by successive steps: the delay from submission to publication, from publication to inclusion in reviews and textbooks, and from there to implementation. It is a composite of averages taken from different studies, which is why Morris, Wooding and Grant, in their 2011 review, found that estimates of time lags in the literature varied enormously and depended on how each study defined its start and end points. Their review is worth dwelling on because it demonstrates the value of the joint-by-joint view. Some lags are between stages, as when a positive trial result waits years for a guideline to incorporate it. Some are within stages, as when a trial takes years to recruit. Some are parallel, as when regulatory review and guideline development proceed simultaneously. Speeding up translation requires knowing which of these dominates for a given kind of intervention, and the answer is different for a cancer drug, a surgical technique, a vaccine and a behavioural programme. There is no single number because there is no single path. What the detailed studies do show, consistently, is that loss is at least as important as delay. The Contopoulos-Ioannidis study described in the introduction found that of 101 highly promising discoveries, only one reached wide clinical use in twenty years. That is not a story of slow progress along a path; it is a story of nearly everything falling off. At the other end of the pipeline, Elizabeth McGlynn and colleagues at RAND, in a 2003 study in the New England Journal of Medicine based on medical records and telephone interviews with adults in twelve US metropolitan areas, found that participants received about 55 percent of the care recommended for their conditions. The failure there is not that the evidence was slow to arrive. It had arrived, been summarised, and been turned into quality indicators. It was simply not being acted upon about half the time. One discovery, five joints The value of the handoff view is clearest when a single intervention is followed all the way along the path. The human papillomavirus vaccine is one of the few for which the whole journey, from laboratory discovery to measured population effect, is now documented, and each of its joints taught a different lesson. At T0, the decisive work was the identification of specific papillomavirus types in cervical cancer tissue. Harald zur Hausen's group in Germany reported HPV16 in 1983 and HPV18 in 1984, against a prevailing view that a herpesvirus was the more likely culprit; the work later earned him a share of the 2008 Nobel Prize in Physiology or Medicine. Establishing that a virus was a necessary cause of most cervical cancers turned a disease of uncertain origin into a potentially preventable infection. But a causal virus is not a vaccine. The critical technical step came in the early 1990s, when Ian Frazer and Jian Zhou at the University of Queensland, and independently groups at the US National Cancer Institute and elsewhere, showed that the viral L1 capsid protein could assemble itself into virus-like particles that carried no genetic material yet provoked a strong antibody response. At T1 and T2, the question became whether those particles were safe, immunogenic and protective in people. A problem arose that recurs in many prevention programmes: the outcome that mattered, invasive cancer, takes decades to develop. The trials therefore used high-grade precancerous lesions as their endpoint, which was defensible biologically but meant that the claim to prevent cancer rested, at licensure, on a surrogate. The quadrivalent vaccine was approved in the United States in June 2006 on the strength of trials such as FUTURE II, reported in the New England Journal of Medicine in 2007. At T3, the programme met a different kind of obstacle. The vaccine worked best if given before sexual debut, which meant vaccinating children and young adolescents against a sexually transmitted infection, an idea that provoked resistance in some countries. Delivery systems mattered enormously: Australia, which began a publicly funded school-based programme in 2007, achieved high coverage quickly, while countries relying on opportunistic vaccination in primary care lagged for years. The same vaccine, with the same efficacy, reached very different fractions of its intended population depending on how it was delivered. At T4, the long-awaited population evidence eventually arrived. A Swedish registry study by Jiayao Lei and colleagues in the New England Journal of Medicine in 2020, covering more than 1.6 million girls and women, found that those vaccinated before the age of 17 had an incidence of invasive cervical cancer close to 90 percent lower than those who were not vaccinated. Scottish data published in 2024 found no cases of invasive cervical cancer among women who had been fully immunised at 12 or 13 in the routine programme. These findings in turn fed policy: the World Health Organization launched a global strategy to eliminate cervical cancer as a public health problem in 2020, and in 2022 its advisory group concluded that a single dose could offer protection comparable to two or three, a change that dramatically lowers the cost of reaching low-income countries. Each joint in this story required a different kind of expertise and was owned by a different set of institutions. At each one, the intervention could have stalled, and in some countries it did. None of the difficulties at T3 or T4 could have been solved by better science at T0. That is the practical meaning of saying that translation fails at the joints. Two directions of travel One further limitation of the linear picture needs to be stated at the outset, because it will recur throughout the book. Knowledge does not flow only from bench to bedside. It also flows back. Some of the most consequential discoveries in medicine began with an observation in patients that sent researchers back to the laboratory. The recognition that a particular chromosomal abnormality was present in the white cells of patients with chronic myeloid leukaemia, first described by Peter Nowell and David Hungerford in Philadelphia in 1960, began at the microscope with patient samples. It took decades of laboratory work to identify the fusion gene responsible and the abnormal enzyme it produced, and then a targeted drug, before imatinib came back to patients at the end of the 1990s. Clinical trials that fail are another source of reverse flow: a well-designed negative trial can reveal that the mechanism studied in animals is not the one operating in humans, and send the question back to T0. Population surveillance at T4 can reveal an unexpected safety signal, a disparity in benefit, or an unanticipated effect that generates new hypotheses for the laboratory. This is why NCATS removed the arrows. The practical significance is that a good translational system does not merely push discoveries forward. It also has channels, and people, whose job is to carry questions backwards, and it treats the information in failures as seriously as the information in successes. Chapter 8 returns to this in detail. What the framework is for A framework of this kind earns its keep if it helps people make better decisions. For a laboratory scientist, the T-framework is a reminder that a discovery's value depends on questions that will be asked at later stages, and that a study designed with those questions in mind is worth more than one designed only to be published. For a clinical investigator, it is a reminder that the trial population, the endpoints and the form of the intervention all determine whether the result will be usable in practice. For a funder, it is a way of seeing where the pipeline is thinnest: historically, far more money has gone to T0 than to T3 and T4 combined, although the precise balance is hard to measure because funders classify their portfolios differently. For a policymaker, it is a way of recognising that evidence of efficacy is not evidence of population effect, and that the absence of population evidence may reflect the absence of anyone paid to generate it. The chapters that follow take the joints in order. The first, between discovery and the decision to test in humans, is where the greatest number of candidates are lost and where the scientific problems are, in some ways, the most fundamental. It is also the joint where the incentives of the sending stage, academic discovery science, are least aligned with the needs of the receiving one. Chapter 2. The First Joint: Why Promising Discoveries Fail to Travel In 2012, C. Glenn Begley, who had spent a decade as head of global cancer research at the biotechnology company Amgen, and Lee Ellis of the MD Anderson Cancer Center published a short comment in Nature that has been cited thousands of times since. Over the preceding decade, Begley's team had attempted to confirm the findings of 53 papers they considered landmark studies in preclinical cancer research, before committing resources to drug programmes built on them. They were able to confirm the scientific findings in only six, about eleven percent. A year earlier, Florian Prinz and colleagues at Bayer had reported in Nature Reviews Drug Discovery that in-house attempts to reproduce published data on potential drug targets had matched the published results in only about a fifth to a quarter of projects. These industry reports had limitations. The companies did not publish which papers they had tested or exactly how, so the claims themselves could not be checked, which is an irony the critics did not fail to note. But they were soon joined by more transparent evidence. The Reproducibility Project: Cancer Biology, a collaboration between the Center for Open Science and Science Exchange, set out to repeat key experiments from high-profile cancer biology papers published between 2010 and 2012. It reported its final results in eLife in 2021. Of the 193 experiments originally selected, the team was able to complete only 50, from 23 papers, in large part because the original papers did not report enough methodological detail and the original authors could not or would not supply it. Among the experiments that were completed, the replication effect sizes were, at the median, about 85 percent smaller than those originally reported. For translational science, these findings describe the condition of the first joint. When a company, a funder or an academic team decides to take a discovery towards human testing, it is relying on a body of preclinical evidence that is, on this evidence, often considerably weaker than it appears. The decision to move to T1 is expensive, and it is taken on the basis of information that the sending stage was not designed to make reliable. Where the weakness comes from It would be comforting to attribute the problem to misconduct, but fraud accounts for a small part of it. Most of the weakness has ordinary causes that operate on honest researchers. The first is design. Preclinical experiments, especially in animals, have historically been small, often unrandomised, and rarely blinded. A series of systematic reviews led by Malcolm Macleod, Emily Sena and colleagues in the CAMARADES collaboration, looking across animal studies of stroke and other conditions, found that studies which did not report randomisation or blinded outcome assessment tended to report larger treatment effects than those which did. That pattern is exactly what one would expect if unconscious bias in allocating animals or scoring outcomes inflates apparent effects. The second is publication bias. Sena and colleagues, in a 2010 paper in PLoS Biology, analysed data from animal studies of stroke and estimated that publication bias alone accounted for roughly a third of the apparent efficacy reported in that literature. Studies with null results are less likely to be written up, less likely to be accepted, and less likely to be cited, so the visible evidence is a selected sample of the more favourable results. The third is analytic flexibility. When an experiment can be analysed in several defensible ways, and the researcher chooses among them after seeing the data, the chance of finding a statistically significant result by chance rises well above the nominal five percent. John Ioannidis's much-discussed 2005 essay in PLoS Medicine, "Why most published research findings are false," made the general argument that small studies, small effects, flexible designs and fields with many competing teams all reduce the probability that a published positive finding is true. Preclinical biology has historically combined all of these. The fourth is the reward structure. A laboratory scientist is rewarded for novelty. A striking result in a high-impact journal can secure grants, promotions and a career. A careful confirmation of someone else's finding, or a well-conducted negative result, rarely does. The sending stage at the first joint is therefore systematically incentivised to produce exactly the kind of evidence the receiving stage finds least reliable: novel, dramatic, underpowered and unconfirmed. Leonard Freedman, Iain Cockburn and Timothy Simcoe attempted to put a price on the consequences. Their 2015 paper in PLoS Biology, drawing on published estimates of irreproducibility rates, suggested that about half of US preclinical research might be irreproducible and that the cost of that irreproducible research was on the order of 28 billion dollars a year. The authors were explicit that the estimate was rough, and the underlying irreproducibility rates are themselves contested. The order of magnitude, however, gives a sense of what is at stake at this joint. The validity problem Even perfectly reproducible preclinical findings can fail to translate, because a reproducible result in a model is not the same as a true result in humans. This is the problem of external or predictive validity, and it is at its most acute with animal models of complex human diseases. Stroke provides the most studied example. Victoria O'Collins and colleagues, in a 2006 systematic review in the Annals of Neurology, catalogued 1,026 experimental treatments for acute ischaemic stroke that had been tested in animals. Of these, 114 had been tested in patients. Only thrombolysis with tissue plasminogen activator had established itself as effective, and aspirin had a modest role. One compound, the free-radical trapping agent NXY-059, became a cautionary tale: it met the field's own recommended criteria for preclinical evidence, produced a marginally positive result in the SAINT I trial, and then showed no benefit in the larger SAINT II trial reported in the New England Journal of Medicine in 2007. Subsequent analyses of its preclinical record found that the animal studies had been small, had often used young, otherwise healthy animals rather than the elderly, hypertensive and diabetic population that suffers most strokes, and had often given treatment soon after the stroke rather than hours later as happens in clinical practice. Sepsis is another. In 2013, Junhee Seok and colleagues, in a large collaborative study in the Proceedings of the National Academy of Sciences, compared genomic responses to trauma, burns and endotoxin in humans and in the mouse models commonly used to study them, and reported that the mouse responses correlated poorly with the human ones. The paper was widely read as an indictment of mouse models of inflammation. A reanalysis of the same data by Keizo Takao and Tsuyoshi Miyakawa, published in the same journal in 2015, reached the opposite conclusion, finding substantial correspondence when the analysis focused on genes that changed significantly in both species. The dispute is instructive in itself. Whether a model is "valid" depends on which features of the disease one needs it to reproduce, and that can only be judged in relation to the specific therapeutic question being asked. Alzheimer's disease offers perhaps the starkest record. Jeffrey Cummings, Travis Morstorf and Kate Zhong, in a 2014 analysis in Alzheimer's Research & Therapy, examined drug-development programmes registered between 2002 and 2012 and calculated a failure rate of 99.6 percent. Many candidates had cleared amyloid or improved cognition in transgenic mice that carried human familial mutations, a form of the disease accounting for a small minority of patients. The eventual approval of antibodies such as lecanemab, which produced modest slowing of decline in large trials, came after two decades of failures, and the debate over how much benefit those drugs provide is still active. The general lesson is that a model's validity has several components that are easy to conflate. Face validity is whether the model looks like the disease. Construct validity is whether it is produced by the same causal mechanism. Predictive validity is whether interventions that work in the model work in patients. A model can have high face validity and poor predictive validity, as some stroke models appear to, or can be useful for one question and useless for another. The first joint leaks badly when models are chosen for convenience and familiarity rather than for their demonstrated ability to predict the human outcome that the next stage will measure. The valley of death To the scientific problems at the first joint must be added an economic one. A discovery made in an academic laboratory is typically far from being a product. Before it can be tested in humans, a candidate drug must be optimised for potency and selectivity, formulated, manufactured to a standard fit for human use, and put through a package of safety pharmacology and toxicology studies designed to satisfy regulators. These activities are expensive, unglamorous, and poorly suited to academic funding, which rewards hypotheses and publications rather than development. Yet they are too risky, at this early stage, to attract most commercial investment. The resulting gap has been called the valley of death since at least the early 2000s, and Declan Butler's 2008 feature in Nature, "Translational research: crossing the valley of death," helped to fix the phrase. Many public programmes since have been designed specifically to bridge it. In the United States, NCATS runs programmes that provide academic investigators with access to industrial-style drug-development expertise, and its Therapeutics for Rare and Neglected Diseases programme was created to advance candidates that no company was likely to pursue. Several universities have established their own drug-discovery units. Venture philanthropy, in which patient foundations fund early development directly, has also played an important role, most famously when the Cystic Fibrosis Foundation invested in the research that led to the drug ivacaftor. The valley of death is not only about money. It is also about knowledge. A laboratory scientist who has discovered a promising target may not know what regulators will require before a first-in-human study, what properties will make a molecule a viable drug, or what the clinical community needs to see before it will enrol patients. The bridging programmes that have worked best supply that knowledge as much as they supply money. Designing the first joint to hold Over the past fifteen years, the response to these problems has taken several forms, and they share a common logic: they make the sending stage accountable for what the receiving stage needs. The first is reporting standards. The ARRIVE guidelines (Animal Research: Reporting of In Vivo Experiments), first published by Carol Kilkenny and colleagues in PLoS Biology in 2010 and revised as ARRIVE 2.0 by Nathalie Percie du Sert and colleagues in 2020, set out what information an animal study needs to report for readers to judge its reliability, including sample-size calculation, randomisation, blinding and the handling of excluded animals. A 2012 paper in Nature by Story Landis and colleagues, arising from a US National Institute of Neurological Disorders and Stroke workshop, called for a core set of transparent-reporting standards along similar lines. Journals and funders have adopted these standards unevenly, and adherence has lagged endorsement, but the direction is clear. The second is funder requirements. From 2016, the NIH began requiring grant applicants to address the rigour of the prior research on which their proposals rested, the rigour of their own designs, the consideration of sex as a biological variable, and the authentication of key resources such as cell lines and antibodies. The sex-as-a-biological-variable policy, announced by Janine Clayton and Francis Collins in Nature in 2014, responded to evidence that preclinical research had relied disproportionately on male animals and cells, so that effects in females were often simply unknown when a product entered human testing. The third is preregistration and confirmatory design. Borrowing from clinical trials, some preclinical researchers now register their hypotheses and analysis plans before collecting data, which removes the analytic flexibility that inflates effect sizes. A more ambitious approach is the multicentre preclinical randomised trial, in which a promising intervention is tested in parallel in several independent laboratories using a common protocol. A 2015 study by Gemma Llovera and colleagues in Science Translational Medicine, which tested an anti-CD49d antibody in experimental stroke across six European centres, showed both that such trials are feasible and that they can overturn findings from single laboratories: the effect was seen in one stroke model but not in another. The fourth is a change in the models themselves. Human-derived systems, including organoids grown from patient stem cells, microphysiological "organ-on-a-chip" devices, and computational models, promise in some applications to predict human responses better than animals do. Regulators have begun to accommodate them. The FDA Modernization Act 2.0, signed into US law in December 2022, removed the statutory requirement that a drug be tested in animals before human trials, allowing sponsors to use other methods where they are adequate. In April 2025, the FDA announced a plan to phase out animal testing requirements for monoclonal antibodies and other drugs, in favour of what it called new approach methodologies. How quickly this changes practice remains to be seen, and for many questions, especially those involving the whole-body interaction of a drug with multiple organs over time, no current alternative fully substitutes for an animal study. One further development deserves separate mention because it reverses the usual direction of evidence at this joint. Instead of starting with a mechanism in a model and hoping it applies to people, researchers increasingly start with evidence from people. Large genetic studies can show that individuals who naturally carry variants disabling a particular gene have lower rates of a disease, which is a kind of natural experiment on what would happen if a drug inhibited that gene's product. The development of PCSK9 inhibitors for lowering cholesterol followed this path: families with gain-of-function mutations in PCSK9 had very high cholesterol, and people with loss-of-function variants had low cholesterol and fewer heart attacks, before any drug had been tested. Matthew Nelson and colleagues at GlaxoSmithKline, in a 2015 analysis in Nature Genetics, found that drug targets with this kind of human genetic support were roughly twice as likely to succeed in development as those without it, and a later analysis by Eric Minikel and colleagues in Nature in 2024 reached a similar conclusion with more data. Human genetic evidence does not replace laboratory work, but it answers, before the first joint is crossed, a question that animal models answer poorly: does modulating this target matter in humans at all? Finally, some of the most important changes are in how decisions are made at the joint. Several pharmaceutical companies now routinely attempt to reproduce key academic findings internally before investing in a programme. AstraZeneca, after a review of its own pipeline published by David Cook and colleagues in Nature Reviews Drug Discovery in 2014, adopted what it called the "5R" framework, asking whether a project had the right target, right tissue, right safety, right patient and right commercial potential, and reported that its later success rates improved. The right patient criterion is especially significant: it asks, at the earliest stage, which people the drug will eventually be tested in, and whether the preclinical evidence speaks to them. That is exactly the kind of downstream question that the first joint has historically neglected. What the first joint hands forward When the first joint works, what passes across it is not merely a molecule and a hypothesis. It is a candidate with reproducible evidence of activity in models chosen for their relevance, an understanding of how it is absorbed, distributed and eliminated, a toxicology package that identifies the organs most at risk, and a biomarker or other measure that can show, in the first human studies, whether the drug is engaging its target. It is also a clear statement of uncertainty: what the preclinical evidence does not show, and therefore what the first human studies will need to find out. Most candidates, even after all this, will fail in humans. But the goal at the first joint is not to eliminate failure; it is to make failure informative and to make it happen early, when it is cheap. A candidate that fails in Phase I because it does not engage its target in humans has taught the field something. A candidate that fails in Phase III because the animal effect was an artefact of unblinded scoring has wasted years and, more seriously, exposed patients to risk for no scientific return. The next chapter follows the candidate across the joint, into the first human beings who will receive it. Chapter 3. Crossing Into Humans: The Logic and Risk of Phase I On the morning of 13 March 2006, eight healthy young men were dosed in a private clinical research unit at Northwick Park Hospital in north-west London. They were the first human beings to receive TGN1412, a monoclonal antibody designed to activate a subset of immune cells by binding to a receptor called CD28. Six received the drug and two received a placebo. The doses were given at intervals of about ten minutes. Within ninety minutes, the six who had received the active drug began to experience severe headache, back pain and fever. Within hours, they had developed what Ganesh Suntharalingam and colleagues, reporting the case in the New England Journal of Medicine later that year, described as a cytokine storm: a massive release of inflammatory signalling molecules leading to multi-organ failure. All six were admitted to intensive care. All survived, but some suffered lasting harm. The dose they had received, 0.1 milligrams per kilogram of body weight, was about five hundred times lower than the highest dose that had caused no adverse effects in cynomolgus monkeys. By the conventional rules for selecting a first human dose, it was cautious. The problem was that the conventional rules assumed that toxicity in animals would predict toxicity in humans, and for this drug, acting on this target, it did not. Subsequent investigation suggested that the immune cells most responsible for the reaction in humans did not express CD28 in the same way in the monkeys used for testing. TGN1412 is the single most influential event in the modern history of first-in-human research, and it illustrates precisely what Phase I is for. It is the stage at which the entire preclinical case is put to its first real test, and where the gap between models and humans becomes, for the first time, directly visible in a living person. What Phase I is trying to learn The classical purpose of a Phase I study is to establish the safety and tolerability of a new intervention in humans, to characterise how the body handles it (its pharmacokinetics) and what it does to the body (its pharmacodynamics), and to identify a dose or range of doses for further study. For most drugs outside oncology, these studies are conducted in healthy volunteers, usually in specialised units, with single ascending doses given to successive small cohorts, followed by multiple ascending doses given over several days. In oncology, the logic is different. Cancer drugs have historically been cytotoxic, damaging healthy cells as well as tumours, and it would be unethical to expose healthy people to them. Oncology Phase I trials are therefore conducted in patients with advanced cancer who have usually exhausted standard options. That changes the ethical calculus: participants may hope for benefit, and the studies have a therapeutic as well as a scientific character. It also changes the design logic, because for cytotoxic drugs the working assumption was that more drug meant more effect, so the goal was to find the highest dose patients could tolerate, known as the maximum tolerated dose. The regulatory scaffolding for Phase I is broadly similar across major jurisdictions. In the United States, a sponsor must submit an Investigational New Drug application to the FDA, including preclinical pharmacology and toxicology data, manufacturing information and a clinical protocol, and may proceed after thirty days unless the agency places the study on hold. In the European Union, clinical trials are now authorised through a single application under the Clinical Trials Regulation, which became applicable in January 2022. In both systems, an independent ethics committee or institutional review board must also approve the protocol, and participants must give informed consent. Choosing the first dose The most consequential single decision in a first-in-human study is the starting dose. Too low, and many cohorts of volunteers will be exposed to a drug at doses that cannot possibly do anything, wasting time and money and raising its own ethical questions. Too high, and the first cohort may be harmed. The traditional approach, codified by the FDA in a 2005 guidance on estimating the maximum safe starting dose in healthy volunteers, begins with the no observed adverse effect level (NOAEL) in the most appropriate animal species, converts it into a human equivalent dose using scaling factors based on body surface area, and then divides by a safety factor, by default ten, to give a maximum recommended starting dose. This approach rests on toxicity: it asks how much drug can be given before something goes wrong in animals. After TGN1412, the UK government convened an Expert Scientific Group, chaired by Gordon Duff, whose report in December 2006 recommended that for high-risk agents, particularly those acting on the immune system with a novel mechanism, the starting dose should instead be based on the minimal anticipated biological effect level, or MABEL. This approach asks not how much drug causes toxicity but how much drug is needed to produce any pharmacological effect at all, using all available data on how the drug binds to its target and what fraction of receptors it occupies at a given concentration. For TGN1412, a MABEL-based starting dose would have been very much lower than the one used. The European Medicines Agency issued a guideline on first-in-human trials for potential high-risk medicinal products in 2007, incorporating MABEL, and substantially revised it in 2017. The other lessons of TGN1412 concerned the conduct of the study. All six active participants had been dosed before any had shown symptoms, because the dosing interval was shorter than the time the reaction took to develop. The revised guidance now emphasises sentinel dosing, in which one or two participants receive the drug first and are observed for an appropriate period before the rest of the cohort is dosed, and dosing intervals chosen on the basis of the drug's expected pharmacology. It also emphasises that studies of high-risk agents be conducted in units with immediate access to intensive care. The second warning: BIA 10-2474 In January 2016, a Phase I trial in Rennes, France, of BIA 10-2474, a drug that inhibits the enzyme fatty acid amide hydrolase, was stopped after one participant was declared brain dead and several others were hospitalised with neurological damage. The affected participants were in a multiple-ascending-dose cohort receiving 50 milligrams a day, the highest dose tested. Unlike TGN1412, the drug had not caused any comparable effect in animals at the doses tested, and other drugs in the same class had been given to humans without serious harm. Investigations by the French medicines agency and an expert committee identified several contributing concerns: the steep escalation between the previous cohort and the one in which harm occurred, doses far above those needed to fully inhibit the target enzyme, and the possibility that the drug had off-target effects on other enzymes in the brain. The episode reinforced a lesson that TGN1412 had already taught. The data that matter for dose escalation are not just whether any adverse effects have been seen, but whether each further increase in dose is actually needed. If the drug has already fully engaged its target at a lower dose, going higher adds risk without adding any therapeutic information. The 2017 revision of the EMA guideline, which was prepared with the Rennes events in mind, emphasises integrating pharmacokinetic, pharmacodynamic and target-engagement data into escalation decisions. Rules for climbing the dose ladder In oncology, where Phase I studies are conducted in patients and the aim has traditionally been to find the maximum tolerated dose, the design of dose escalation has been a subject of intense statistical debate. The approaches in widest use differ in how they decide whether to increase, hold or reduce the dose for the next cohort, and in how efficiently they identify the right dose, as Table 2 summarises. Table 2. Common dose-escalation designs in early oncology trials. Design How escalation is decided Main strength Main weakness 3+3 Fixed rules based on toxicities in cohorts of three Simple, familiar, needs no statistician at the bedside Often misidentifies the target dose; treats many patients at low doses Continual reassessment method (CRM) A statistical model of dose and toxicity, updated after each cohort Uses all accumulated data; identifies the target dose more accurately Requires modelling expertise; perceived as opaque Bayesian optimal interval (BOIN) Compares the observed toxicity rate at the current dose with pre-set boundaries Nearly as accurate as model-based designs; simple to run Still focused on toxicity rather than benefit The 3+3 design, which dates in essence from the 1970s and 1980s, treats three patients at a dose; if none experiences a dose-limiting toxicity, the next three receive a higher dose; if one does, three more are added at the same dose; if two or more do, escalation stops and the dose below is usually declared the maximum tolerated dose. It is transparent and easy to run. It is also, according to a large body of simulation studies, poor at identifying the dose it aims to find, and it treats a large fraction of patients at doses well below any likely to be effective. The continual reassessment method, introduced by John O'Quigley, Margaret Pepe and Lloyd Fisher in Biometrics in 1990, fits a statistical model relating dose to the probability of toxicity and updates it after each patient or cohort, recommending the dose whose estimated toxicity is closest to a target rate. It uses information more efficiently and more patients are treated near the target dose. Its adoption was slowed by concerns, some justified in its earliest versions, that the model could escalate too aggressively, and by the practical need for statistical support throughout the trial. The Bayesian optimal interval design, described by Suyu Liu and Ying Yuan in 2015, occupies a middle ground: it uses pre-calculated decision boundaries that can be written down in a table at the start of the trial, so that it is as easy to run as the 3+3 design while performing much closer to model-based methods. When the maximum is not the optimum All three designs share an assumption inherited from cytotoxic chemotherapy: that the right dose is the highest dose patients can tolerate. For many modern cancer drugs, that assumption is wrong. Targeted therapies that inhibit a specific enzyme may achieve full target inhibition at doses well below those that cause intolerable side effects. Immunotherapies may have flat dose-response relationships across a wide range. For such drugs, pushing to the maximum tolerated dose adds toxicity without adding benefit, and patients who take the drug for months or years may suffer side effects that lead them to reduce doses or stop treatment altogether. The FDA's Oncology Center of Excellence launched Project Optimus in 2021 to address this, and in 2024 published final guidance on optimising the dosage of oncology drugs. It asks sponsors to compare multiple doses, often in randomised fashion, before committing to a registrational trial, and to characterise the relationships between dose, exposure, activity and safety. A frequently cited illustration is sotorasib, a drug targeting a mutated form of the KRAS protein, which received accelerated approval in May 2021 at a dose of 960 milligrams daily. As a condition of approval, the FDA required a post-marketing trial comparing that dose with a quarter of it, reflecting uncertainty about whether the lower dose would be as effective with fewer side effects. This is a case where the handoff at the second joint, from Phase I to later development, had been designed around the wrong question. A Phase I design optimised to find the highest tolerable dose was handing forward a dose that later stages, and patients, did not need. The correction required changing what Phase I was asked to deliver. The ethics of being first The ethical issues in Phase I differ between healthy-volunteer studies and patient studies, but both concern the relationship between risk and benefit when the participant is the first human being to take a drug. For healthy volunteers, there is no prospect of direct benefit, and participation is typically paid. The concern is whether payment induces people to accept risks they would otherwise refuse, and whether the population of habitual paid volunteers, often economically marginal, bears a disproportionate share of the burden of drug development. The empirical record on harm is, on the whole, reassuring. Ezekiel Emanuel and colleagues, analysing data from healthy-volunteer Phase I studies conducted by a large pharmaceutical company in a 2015 paper in the BMJ, found that serious adverse events were rare and most adverse events were mild. The rarity of disasters such as TGN1412 and BIA 10-2474 is part of what makes them so instructive. For patients in oncology Phase I trials, the concern is different: whether patients understand that the primary purpose of the trial is to find a dose, not to treat them, and whether the prospect of benefit is realistic. Elizabeth Horstmann and colleagues, in a 2005 analysis in the New England Journal of Medicine of 460 Phase I oncology trials conducted between 1991 and 2002, found an overall response rate of 10.6 percent and a rate of death attributed to toxicity of 0.49 percent, with response rates higher in trials that included an established anticancer agent. More recent analyses have reported higher response rates in the era of targeted therapy. These data have been used to argue that Phase I participation offers more potential benefit than had often been assumed, and that describing these trials as purely non-therapeutic in consent discussions is inaccurate. They have also been used to argue that the true chance of benefit varies so widely by drug and setting that honest consent requires trial-specific information, not general reassurance. A different response to the ethics of first-in-human exposure is to reduce what needs to be learned at full dose. The FDA's 2006 guidance on exploratory Investigational New Drug studies allowed very limited early human studies, sometimes called Phase 0, in which subtherapeutic microdoses are given to a small number of people to learn about how a drug is distributed or whether it reaches its target. Such studies cannot establish safety or efficacy, but they can end a programme early, at low risk, if the drug behaves in humans very differently from how it behaved in animals. What Phase I should hand forward The standard output of Phase I, a recommended dose and a list of observed adverse events, is necessary but not sufficient. The next stage needs to know not just what dose can be tolerated but whether the drug reached its target, whether it engaged it, and whether that engagement produced the expected biological effect. These three questions, sometimes called the three pillars of survival in drug development, were articulated in a 2012 analysis by Paul Morgan and colleagues at Pfizer in Drug Discovery Today. Reviewing the company's Phase II programmes, they found that many failures were in programmes where it had never been established that the drug had reached and engaged its target at the doses tested. Without that information, a negative Phase II result cannot distinguish between a drug that does not work and a drug that was never given a fair chance. A Phase I study designed with the next stage in view therefore measures pharmacodynamic biomarkers alongside safety, explores more than one candidate dose, and characterises how exposure varies between people. It hands forward not just a number but an understanding of the dose-exposure-response relationship on which every later decision will depend. That understanding becomes indispensable in the stage the next chapter examines, where most candidates that survive Phase I will fail. Hashtags: #TranslationalScienceFundamentals #TranslationalScience #BenchToBedside #TranslationalResearch #T0ToT4 #PreclinicalResearch #FirstInHumanStudies #PhaseITrials #ClinicalTrials #ImplementationScience #DisseminationScience #PopulationHealth #Reproducibility #PredictiveValidity #ValleyOfDeath #DrugDevelopment #DoseOptimization #Pharmacokinetics #Pharmacodynamics #TargetEngagement #EvidenceToPractice #ImplementationResearch #TranslationalMedicine #ClinicalResearch #FutureOfTranslationalScience
- Agent-Based Modeling in Social Epidemiology (Simulating Contagion and Human Behavior)
Download the Book (PDF): Introduction In the spring of 2020, public health agencies in dozens of countries faced a question that no amount of surveillance data could answer on its own. If schools closed, if workplaces sent people home, if the elderly were asked to shield, how many infections would follow, and where? The data describing the epidemic were already weeks out of date by the time they were reported, and the interventions under discussion had never been tried at scale on this pathogen. Decisions had to be made about a future that did not yet exist, in a population whose behavior was changing week by week in response to the very epidemic it was trying to escape. The tools that governments reached for were simulations. Some were compartmental models, the century-old descendants of the equations William Kermack and Anderson McKendrick published in 1927, which divide a population into the susceptible, the infected and the recovered and track the flows between them. Others were individual-based simulations of whole national populations, in which millions of synthetic people lived in households, attended schools and workplaces, and passed the virus along the contacts those settings created. The most famous of these, the model behind the Imperial College COVID-19 Response Team's Report 9 of 16 March 2020, descended from work on pandemic influenza that Neil Ferguson and colleagues had published in Nature in 2005 and 2006. Covasim, developed at the Institute for Disease Modeling, was released as open-source Python and adapted by teams on several continents within months. These models attracted intense public scrutiny, much of it confused. Critics attacked their code, their assumptions, and their forecasts, often without distinguishing among the three. Defenders sometimes spoke as though a simulation could be judged by whether its headline number came true, when the entire point of a scenario model is that its projections trigger actions that change the outcome. What almost nobody discussed in public was the question that matters most to a social epidemiologist: how the models represented people. Who met whom, and why? How did individuals decide to stay home, wear a mask, accept a vaccine, or ignore the advice? Why did infection and death fall so unevenly across neighborhoods, occupations, and ethnic groups? The epidemic was, from its first weeks, a social phenomenon as much as a biological one, and the models that were best at representing its biology were often the weakest at representing its society. This booklet is about building simulations that take both seriously. Its subject is agent-based modeling: the practice of representing a population as a collection of individual actors, each with its own attributes, location, relationships, and rules for behavior, and letting population-level patterns emerge from their interactions. Its domain is social epidemiology, the branch of epidemiology concerned with how social structure, from income and housing to networks and norms, shapes the distribution of health and disease. The two belong together because both are concerned with the same problem. Social epidemiology asks how the organization of society produces patterns of disease that no individual chose. Agent-based modeling is a method for asking exactly that question in a form a computer can answer. The argument of this booklet The controlling idea is simple to state and demanding to practice. An agent-based model earns its place in social epidemiology when, and only when, the question depends on feedback among three things that aggregate methods average away: the heterogeneity of individuals, the structure of their contacts, and the way their behavior responds to what is happening around them. When those feedbacks matter, a well-built agent-based model can show how a population-level pattern is generated, test interventions that cannot ethically or practically be tried, and expose which unknowns actually drive the answer. When they do not matter, an agent-based model is an expensive way to reproduce what a differential equation would have told you in an afternoon, with more places for error to hide. It follows that the discipline of the method lies less in programming than in judgment. The hard decisions are about which question the model is for, which mechanisms it must contain, what data can constrain it, and how its uncertainty will be reported. Code matters, and later chapters treat it concretely, but a model written flawlessly in NetLogo or Python around a vague question is still a vague model. Researchers who come to agent-based modeling from statistics often expect the difficulty to be technical. It is mostly conceptual: learning to state a mechanism precisely enough that a machine can execute it, and honestly enough that a reader can dispute it. What the chapters do Chapter 1 sets out what agent-based models add to the compartmental tradition and, just as important, what they cost. Chapter 2 turns to design: how to begin from a question rather than from a platform, how to decide what a model should leave out, and how the ODD protocol gives a model a description that others can read, criticize, and reproduce. Chapter 3 addresses the structure of contact, the networks and places through which both pathogens and behaviors travel, and the synthetic populations on which large models rest. Chapter 4 moves inside the host, showing how natural history, infectiousness, and transmission probability are represented at the level of the individual agent, and why stochasticity is not a nuisance but a finding. Chapters 5 and 6 are the heart of the social side of the argument. Chapter 5 is about individual decision-making: how agents perceive risk, weigh costs, and adopt or abandon protective behavior, and how psychological theory can be translated into rules without pretending to more precision than it has. Chapter 6 is about social influence: the difference between simple and complex contagion, the problem of distinguishing influence from homophily, and the ways in which segregation and inequality become embedded in the structure of a simulated society. Chapter 7 is practical, working through how models are actually built in NetLogo and in Python, with attention to the habits that make code trustworthy. Chapter 8 addresses calibration, validation, and sensitivity analysis, the work that separates a model someone can rely on from one that merely runs. Chapter 9 takes up the use of models for causal reasoning and policy, including their special role in questions of equity. Who this is for The reader imagined here is a researcher who knows epidemiology or a neighboring social science, who can read a regression table and has probably written some code, and who wants to understand how to build and judge simulations of contagion and behavior. No prior modeling experience is assumed, but the booklet does not linger on programming basics; where code matters, it describes what the code must do and why. Equations appear rarely and are always explained in words. A word on scope. Agent-based modeling reaches well beyond infectious disease, into obesity, violence, substance use, and the health consequences of urban form, and several of those applications appear here because they illuminate the method. But the organizing thread is contagion in its double sense: the transmission of pathogens between bodies, and the transmission of behaviors, beliefs, and norms between minds. The most interesting problems in the field sit where the two meet, where fear spreads faster than a virus, where a vaccine refusal cluster turns a controlled disease into an outbreak, where a neighborhood's exposure is fixed by the jobs its residents cannot do from home. Those are the problems an agent-based model is uniquely equipped to address, and they are the problems this booklet is written to help its reader model well. Chapter 1. Why Agents? From Compartments to Individuals Every model of an epidemic is a claim about how people meet. The claim may be explicit or buried in an assumption, but it is always there, and it determines much of what the model can say. Understanding agent-based modeling begins with understanding the claim made by the models that preceded it, because the reasons for building agents at all are the places where that older claim breaks down. The compartmental tradition and its bargain In 1927 Kermack and McKendrick published a set of equations describing an epidemic in a closed population. People are sorted into compartments: susceptible (S), infectious (I), and removed (R), later more commonly called recovered. The rate at which new infections occur is proportional to the product of the number of susceptible and the number of infectious people, multiplied by a transmission coefficient usually written as beta. Infectious people leave the infectious compartment at a constant rate, gamma, whose inverse is the mean duration of infectiousness. From these few assumptions follow some of the most important results in public health: that an epidemic grows only if each case, in a fully susceptible population, produces on average more than one further case; that this basic reproduction number, R0, equals beta divided by gamma in the simplest version; and that an epidemic stops before everyone is infected, once the susceptible fraction falls far enough that each case replaces itself less than once. The herd immunity threshold, one minus the reciprocal of R0, comes directly from this logic. The product of S and I is the crux. It encodes the assumption of mass action, borrowed from chemistry: individuals mix like molecules in a well-stirred solution, so that any infectious person is equally likely to meet any susceptible one. Everyone in a compartment is interchangeable. Nobody has a household, a job, a friend, or a neighborhood. Nobody decides anything. This is not a flaw so much as a bargain. By giving up individual detail, the modeler gains analytical tractability, a small number of parameters that can often be estimated from data, and results that generalize. The compartmental framework has been extended enormously since 1927: exposed but not yet infectious compartments (the SEIR model), age structure through contact matrices, waning immunity, vaccination compartments, spatial patches linked by travel. Roy Anderson and Robert May's Infectious Diseases of Humans (1991) remains the classic synthesis of what this tradition can do, and it can do a great deal. A modeler who ignores it and reaches directly for agents is usually making a mistake. The bargain fails in identifiable circumstances. It fails when the heterogeneity that has been averaged away is itself the thing driving the outcome. It fails when the structure of contact, who is connected to whom and how tightly, changes the dynamics in ways that mean contact rates cannot capture. It fails when people's behavior responds to the epidemic, and does so differently depending on their circumstances, their information, and the behavior of those around them. And it fails when the question itself is about individuals or small groups: which households, which clusters, which neighborhoods. Each of these failures can sometimes be patched within the compartmental framework by adding more compartments, but the patches multiply quickly. An SEIR model stratified by five age groups, three income levels, two vaccination states, and two behavioral states already has 180 compartments, and it still assumes mixing is random within and between them according to some matrix the modeler must specify. What an agent-based model is An agent-based model, in the sense used throughout this booklet, has four components. There are agents, discrete entities with their own state: an age, a disease status, a location, a set of relationships, perhaps a belief about risk or a tendency to comply with advice. There is an environment in which agents exist, which may be a spatial grid, a map of real places, a network, or some combination. There are rules governing how agents act and interact, including how they move, whom they contact, how infection passes between them, and how they make decisions. And there is a schedule, the order in which things happen within each time step and across the simulated period. The model runs by applying the rules repeatedly and recording what happens. Nothing about the population-level trajectory, the epidemic curve, the final attack rate, the distribution of cases across groups, is written into the model directly. It emerges from the accumulation of individual events. Joshua Epstein, whose Generative Social Science (2006) gave the approach its most influential methodological statement, put the underlying standard of explanation in a phrase: if you did not grow it, you did not explain it. To explain a pattern generatively is to show a population of agents, following specified rules, that produces the pattern. The approach grew from several roots. Thomas Schelling's models of residential segregation, published in 1971 and popularized in Micromotives and Macrobehavior (1978), showed that mild individual preferences about neighbors could produce stark segregation that no individual wanted. Epstein and Robert Axtell's Sugarscape, described in Growing Artificial Societies (1996), grew trade, migration, inequality, and disease transmission from agents on a landscape of resources. In ecology, where the approach is usually called individual-based modeling, Volker Grimm, Steven Railsback, and colleagues developed much of the methodology for describing, testing, and validating such models. In infectious disease epidemiology, individual-based simulation became central to pandemic preparedness planning in the 2000s, through the influenza models of Ira Longini, Ferguson, and Timothy Germann and colleagues, which represented whole national populations in their households, schools, and workplaces. In social epidemiology specifically, the case for the method was made in a sequence of papers around 2008 to 2012. Amy Auchincloss and Ana Diez Roux argued in the American Journal of Epidemiology in 2008 that dynamic agent models could address place effects on health that conventional multilevel regression treats awkwardly, because regression holds the context fixed while in reality people and places shape each other over time. Sandro Galea, Matthew Riddle, and George Kaplan (2010) argued that the complex, interdependent causation typical of social determinants of health strains the assumptions of standard causal inference and calls for complex systems methods. Abdulrahman El-Sayed and colleagues (2012) reviewed the joint use of social network analysis and agent-based modeling in social epidemiology. The argument in all of these was the same at root: that the phenomena social epidemiologists care about are generated by interacting people in structured environments, and that a method which represents those interactions directly can reveal things that methods treating individuals as independent observations cannot. What agents add Four capacities distinguish agent-based models from their aggregate counterparts, and each corresponds to one of the failures of the compartmental bargain. The first is heterogeneity without combinatorial explosion. Each agent can carry as many attributes as the question requires, in any combination, drawn from real joint distributions. An agent can be a 67-year-old bus driver with diabetes living in a four-person household in a particular census tract, and the model does not need a separate compartment for every such combination; it simply has a list of people. The second is explicit contact structure. Rather than assuming random mixing, an agent-based model can specify exactly who is in contact with whom, through households, workplaces, schools, friendship networks, sexual partnerships, or spatial proximity. This matters because network structure changes epidemic dynamics in ways averaged rates cannot represent. Hazhir Rahmandad and John Sterman made the comparison directly in Management Science in 2008, building an agent-based SEIR model on several network types and comparing it with the equivalent differential equation model. On highly clustered networks such as lattices, epidemic dynamics differed substantially from the aggregate model's predictions; on random networks with some connectivity, the aggregate model did well. Their conclusion was measured: network structure matters when it is strongly clustered or when degree is very unequal, and the aggregate model is often adequate otherwise. That measured conclusion is the right starting point for any modeler. The third capacity is adaptive behavior. Agents can observe their surroundings, update beliefs, and change what they do. An agent can reduce contacts when local prevalence rises, adopt a mask because its friends did, refuse a vaccine because rumors of harm reached it through its network, or return to work when savings run out. Because each agent's decision depends on its own situation and neighborhood, the model can represent behavior that is not only responsive but unevenly responsive, which is precisely what social epidemiology expects. The fourth is stochasticity at the level of events. Every infection in an agent-based model is a discrete random event. Early in an outbreak, when a handful of cases determine whether a chain of transmission takes off or dies out, this matters enormously, and deterministic compartmental models, which have fractional people flowing smoothly between boxes, cannot represent it at all. Stochastic compartmental models can, but only for the averaged population they describe. The differences among model families on these dimensions are summarized in Table 1, which compares deterministic compartmental models, network models, and full agent-based models on the features that most often decide the choice among them. Table 1. Three families of epidemic model compared on the features that usually decide the choice. Feature Compartmental (ODE) Network model Agent-based model Unit represented Population fractions Nodes and edges Individuals with attributes and rules Contact assumption Mass action within groups Fixed or evolving graph Any: places, networks, space, movement Heterogeneity Only by adding compartments Mostly in degree and position Arbitrary combinations of attributes Behavior Usually fixed or imposed Sometimes rewiring rules Adaptive decision rules per agent Stochastic extinction Absent unless made stochastic Present Present Analytic results Often available Some (thresholds, moments) Rarely; relies on simulation Data demand Low Moderate High What agents cost The same capacities that make agent-based models powerful make them costly, and an honest account of the method has to lead with the costs rather than bury them. The most important cost is parameter proliferation. Every rule has parameters, and every attribute has a distribution. A model with realistic contact structure, a detailed natural history, and adaptive behavior can easily have fifty or a hundred parameters, many of which have never been measured. The modeler must choose values for all of them, and the results may depend on choices that no data constrain. This is the core problem that Chapter 8 addresses, and it is not solved by adding more detail; more detail usually makes it worse. The second cost is opacity. Because results are generated by simulation rather than derived, it can be hard to know why a model produced the output it did. A compartmental model's behavior can often be read off its equations; an agent-based model's must be discovered by experimentation on the model itself. Without deliberate effort, including systematic sensitivity analysis and careful documentation, a complex agent-based model becomes a black box to its own authors. The third is computational expense. Stochastic models must be run many times to characterize the distribution of outcomes, and each run of a large model may take minutes or hours. Calibration and sensitivity analysis multiply the number of runs required, sometimes into the hundreds of thousands. This constrains which questions can be explored thoroughly. The fourth is verification risk. The more code there is, the more places there are for errors. Uri Wilensky and William Rand's 2007 attempt to replicate a published agent-based model, reported in the Journal of Artificial Societies and Social Simulation, found that seemingly minor ambiguities in how the original was described, such as the order in which agents were updated, produced materially different results. Later replication efforts have repeatedly found the same thing. A result that depends on an unreported implementation detail is not a result about the world. Choosing the right tool The practical question for any researcher is whether a particular problem calls for agents. A useful discipline is to try to state the question in a form a compartmental model could answer, and see what gets lost. Suppose the question is how much a seasonal influenza vaccination campaign reduces hospitalizations in a city, given known coverage by age. An age-structured SEIR model with a contact matrix can answer that, and has done so in many published analyses. An agent-based model would add little except cost. Now suppose the question is how much the same campaign reduces hospitalizations when vaccine uptake is clustered by neighborhood and social network, because refusal spreads among friends and is concentrated in particular communities. The compartmental model can include a vaccinated compartment, but it cannot easily represent the fact that the unvaccinated are clustered together, and clustering matters: Marcel Salathé and Sebastian Bonhoeffer showed in 2008, in the Journal of the Royal Society Interface, that when vaccine refusers cluster in a network, outbreak probability can rise substantially even at the same overall coverage. Now the question depends on contact structure and on the social process generating uptake. This is agent territory. Or suppose the question is why a policy of free healthy food vouchers did not reduce income disparities in diet as much as expected. The answer might depend on where low-income households live relative to stores, how store locations respond to demand, and how preferences spread through neighborhoods. Auchincloss and colleagues built an agent-based model of exactly this kind of system, reported in the American Journal of Preventive Medicine in 2011, in which households and food stores were placed on a landscape with varying degrees of income segregation. Absent other factors, income differences in diet emerged from the segregation of high-income households and healthy stores from low-income households and unhealthy stores. When both income groups shared a preference for healthy food, low-income diets improved but a disparity remained; only a combination of favorable preferences and relatively cheaper healthy food overcame the gap generated by segregation. No regression on observational data could have produced that finding, because it is about the interaction of multiple mechanisms over time in a spatial structure. It is a finding about how a pattern is generated. The general rule that emerges from such contrasts is this: reach for agents when the answer depends on who interacts with whom, when behavior responds to local conditions, or when the question concerns the distribution of outcomes across a structured population rather than their average. Otherwise, start simpler. Many of the best agent-based studies in epidemiology were built alongside, or after, a simpler compartmental model of the same system, and the comparison between them is often the most informative result. When the agent-based model agrees with the aggregate one, the modeler learns that heterogeneity and structure do not matter much for this question. When it disagrees, the modeler has found exactly the mechanism that deserves attention. That comparison also makes an important point about what agent-based models are for. They are not simply more realistic versions of compartmental models, to be preferred whenever computing power allows. They are instruments for a particular kind of reasoning: reasoning about how individual-level mechanisms, operating in structured populations, generate the patterns epidemiologists observe. Used for that purpose, they can do things no other method can. Used as a default, they tend to generate complicated answers to simple questions and unwarranted confidence about complicated ones. Chapter 2. Starting With the Question: Purpose, Scope, and the ODD Protocol Most failed agent-based models fail before any code is written. They fail because the modeler began with a platform, a dataset, or an enthusiasm for realism, rather than with a question sharp enough to say what the model must contain and what it can safely ignore. This chapter is about the decisions that come first: what the model is for, how much it should include, and how it should be described so that others can evaluate it. Models have purposes, and purposes set standards It is tempting to think of a model as a miniature copy of the world, better the closer it resembles the original. That view leads straight to models that include everything and explain nothing. A more useful view is that a model is a tool built for a job, and that the job determines how the model should be judged. Epstein's 2008 essay "Why Model?" in the Journal of Artificial Societies and Social Simulation listed sixteen reasons to build models other than prediction, including explaining, guiding data collection, illuminating core dynamics, suggesting analogies, discovering new questions, bounding outcomes, and revealing the apparently simple to be complex. Bruce Edmonds and colleagues, writing in the same journal in 2019, sharpened this into a set of distinct purposes, each with its own standard of success. Among those most relevant to epidemiology are prediction, which requires that a model reliably anticipate unknown data; explanation, which requires showing that a plausible mechanism produces an observed outcome; theoretical exposition, which explores the consequences of a set of assumptions without claiming they describe any particular system; and illustration, which communicates an idea through a simple, vivid case. These purposes are frequently confused, and the confusion causes trouble. A model built to illustrate how clustering of vaccine refusal can amplify outbreaks does not need to be calibrated to any real city, and criticizing it for failing to predict measles cases in a particular county misunderstands it. Conversely, a model presented as a forecast for policy needs to be validated against out-of-sample data, and defending it on the grounds that its mechanisms are plausible is not enough. Much of the public argument over COVID-19 models in 2020 involved exactly this confusion: scenario projections, which say what would happen under stated assumptions, were treated as predictions, and then condemned when the assumptions changed because policy and behavior did. For social epidemiology, explanation and scenario analysis are usually the most defensible purposes. The typical question is either of the form "how could this pattern arise?" or "what would happen to this outcome, and its distribution across groups, if this intervention were implemented?" Pure prediction of an epidemic's future course is a legitimate aim but a much harder one, and agent-based models have no special advantage at it. How much to include Once the purpose is clear, the next question is scope. Two opposing philosophies have long competed. One, often summarized as KISS (keep it simple, stupid), holds that models should be as simple as possible, adding complexity only when simplicity demonstrably fails. The other, which Edmonds and Scott Moss labeled KIDS (keep it descriptive, stupid) in a 2005 paper, holds that social models should begin with as much empirically grounded description as possible and simplify only when detail proves unnecessary. Both have merit, and the choice between them depends partly on purpose: a theoretical exposition benefits from simplicity, while a scenario model for a specific city may need descriptive richness to be credible to decision-makers. The most useful practical guide to scope is Grimm and colleagues' pattern-oriented modeling, set out in Science in 2005 and developed in Railsback and Grimm's textbook. Its central idea is that a model should be complex enough to reproduce several independent patterns observed in the real system, and no more complex than that. A single pattern, such as the shape of an epidemic curve, can be reproduced by many different mechanisms, so matching it tells you little. Several patterns at different levels, say the epidemic curve, the age distribution of cases, the proportion of infections occurring in households, and the geographic spread over time, jointly constrain the model far more tightly. A mechanism that reproduces all of them simultaneously is much more likely to be right than one tuned to reproduce any one. Pattern-oriented modeling gives the modeler a principled way to decide what to include. For each candidate mechanism, the question becomes: which observed pattern requires it? If the household secondary attack rate is one of the target patterns, the model needs households. If the socioeconomic gradient in infection is a target, the model needs something that generates it, perhaps occupation-based exposure or household crowding. If no target pattern requires a mechanism, and the question does not turn on it, it can be left out. A few concrete heuristics follow from this approach. Start from the smallest model that could answer the question, and write down, before building it, what the simplest version would predict. This is the baseline against which every addition is judged. Add mechanisms one at a time, and for each, check whether it changes the answer to the question. If it does not, consider removing it; unnecessary mechanisms add parameters and obscure interpretation. Distinguish between mechanisms the question is about and mechanisms that merely need to be present for the model to be plausible. The first deserve careful, theory-based representation. The second can often be represented crudely, as long as sensitivity analysis confirms the answer does not depend on the crude choice. Be suspicious of detail driven by available data rather than by the question. A rich survey dataset makes it tempting to give agents dozens of attributes, but attributes that do not enter any rule do nothing except slow the model down, and attributes that enter rules without theoretical justification introduce unexamined assumptions. A worked example of scoping Consider a research team that wants to understand why influenza vaccination coverage among adults in a mid-sized city varies so sharply across neighborhoods, and how much that variation matters for citywide influenza burden. A naive specification might begin by listing everything known about influenza and vaccination and trying to include it all. A question-first approach begins instead by stating the purpose, which here is scenario analysis with an explanatory component: to estimate how much influenza burden would change if coverage in low-uptake neighborhoods rose to the city average, compared with an equal number of extra doses distributed uniformly. That purpose immediately identifies what the model must contain. It needs neighborhoods, because the question is about their differences. It needs contact patterns within and across neighborhoods, because whether low-coverage neighborhoods are well mixed with the rest of the city determines whether their susceptibility spills over. It needs an age structure, because both vaccination and influenza severity vary strongly by age. It needs a representation of vaccination status that reflects actual neighborhood coverage. What it does not obviously need is a behavioral model of vaccine decision-making, because the scenario sets coverage directly rather than asking how coverage would respond to some intervention. If the team later wants to ask which intervention would raise coverage in low-uptake neighborhoods, a decision model becomes necessary, and that is a different, harder model. Separating the two questions keeps the first model tractable and makes clear what the second would require. The target patterns for validation might include the observed age distribution of influenza hospitalizations, the observed neighborhood variation in influenza-like illness from surveillance, and the ratio of household to community transmission reported in the literature. Each constrains a different part of the model. Only once these are specified does it make sense to choose a platform and write code. The ODD protocol A model that cannot be described cannot be evaluated. For many years agent-based models were described in idiosyncratic ways, some in prose, some in code listings, some barely at all, and replication was correspondingly difficult. In 2006 Grimm and a large group of ecological modelers published in Ecological Modelling a standard format for describing individual-based and agent-based models, which they called ODD, for Overview, Design concepts, and Details. The protocol was updated in 2010 and again in 2020, the latter update published in the Journal of Artificial Societies and Social Simulation. It has become the most widely used description standard across ecology and the social sciences, and increasingly in epidemiology. The current protocol has seven elements grouped under its three headings. The Overview covers the model's purpose and the patterns used to evaluate it; its entities, state variables, and scales, meaning what kinds of agents and environments exist, what attributes each has, and the spatial and temporal resolution and extent; and its process overview and scheduling, meaning what happens in each time step and in what order. The Design concepts element asks the modeler to address a checklist of conceptual issues: the basic principles and theories underlying the model, which outcomes emerge and which are imposed, how agents adapt, what objectives they pursue, whether they learn or predict, what they sense, how they interact, where stochasticity enters, whether collectives such as households exist as entities, and what is observed and recorded. Not every concept applies to every model, and the protocol allows a modeler to say so, but working through the list forces decisions that otherwise remain implicit. The Details cover initialization, meaning the state of the model at the start of a run and whether it varies across runs; input data, meaning any external data that drive the model during a run, such as time series of policy changes; and submodels, meaning the full specification of each process listed in the overview, with its equations, parameters, and justification. Writing an ODD description is often the moment a modeler discovers that the model is not yet fully designed. It is easy to say agents "reduce contacts when they perceive risk." It is much harder to say which contacts, by how much, in response to what information, updated how often, and in what order relative to transmission within a time step. The ODD's insistence on details makes this vagueness visible before it becomes code. The protocol also improves the model itself in less obvious ways. The requirement to specify scheduling, for example, exposes a class of error that has tripped up many published models: whether agents update their states synchronously, all at once based on the previous step's state, or asynchronously, one at a time with each seeing the updates of those who went before. In an epidemic model, this determines whether an agent infected in the current step can infect others in the same step, and the choice can change the growth rate materially. Specifying it forces a decision, and the decision should be justified by the biology and the time step. Several extensions to ODD are worth knowing. ODD+D, proposed by Birgit Müller and colleagues in 2013, adds guidance for describing human decision-making, including the theoretical basis for decision rules, which is especially relevant to Chapter 5. The TRACE framework, developed by Grimm and colleagues, documents the modeling cycle itself: how the model was tested, calibrated, and analyzed. In health economics and outcomes research, the ISPOR-SMDM Modeling Good Research Practices Task Force published guidance in 2012 on dynamic transmission modeling and on model transparency and validation, which complement ODD for models meant to inform policy. Drawing on people who know the system A model of a social system is only as good as its modelers' understanding of that system, and in social epidemiology much of the relevant knowledge lives outside the research team. The people who inject drugs in a particular city know how syringes actually circulate; school nurses know which contacts dominate a school day; community health workers know why a neighborhood distrusts a vaccination clinic. A model built without that knowledge will encode the researchers' guesses about mechanisms, and those guesses are often wrong in ways that no amount of calibration can repair. The HIV model built by Brandon Marshall, Sandro Galea, Samuel Friedman, and colleagues for the New York metropolitan area, published in PLoS ONE in 2012, illustrates the alternative. The team was deliberately multidisciplinary, drawing epidemiologists, sociologists, geographers, and mathematicians together, and its conceptual framework was built from prior ethnographic as well as epidemiological research on drug-using populations. Agents represented people who inject drugs, people who use drugs without injecting, and people who use no drugs, interacting within sexual and injection risk networks and encountering syringe exchange programs, substance use treatment, HIV testing, and antiretroviral therapy. The model was calibrated against observed trajectories of HIV prevalence and incidence and reproduced the decline in HIV prevalence among people who inject drugs in New York City between 1992 and 2002. Its value for exploring combination prevention depended on the conceptual framework being right about how risk networks and service contact actually worked, which is knowledge the ethnographic record supplied. Formal participatory approaches go further. In group model building and companion modeling, stakeholders help draw the causal structure of the system, critique early versions of the model, and interpret its results. These methods come largely from system dynamics and natural resource management, and they are slower and messier than a research team working alone. But they catch mistaken assumptions early, they surface mechanisms that the literature has not described, and they produce models that the people expected to act on them understand and trust. For models intended to inform local policy, particularly in communities that have reason to be wary of research done about them rather than with them, that trust is not a side benefit. It determines whether the model is used at all. Even without a formal participatory process, a modeler can adopt the habit of showing the conceptual model, a plain-language account of agents, their attributes, and the causal pathways among them, to people with direct knowledge of the system before writing the details. The question to ask them is not whether the model is realistic, since every model is unrealistic in countless ways, but whether anything it leaves out would change the answer to the question it is meant to address. Common design failures Several failures recur often enough to name. The first is the everything model, built to answer any question about a system and therefore well suited to none. It typically has hundreds of parameters, takes hours to run, and produces outputs its own authors struggle to explain. The remedy is to build separate, smaller models for separate questions, sharing components where that is convenient. The second is the unstated mechanism: a rule included because it seemed reasonable, never justified by theory or data, that turns out to drive the results. Behavioral rules are especially prone to this. An agent that reduces contacts by half when it knows one infected person is making a strong behavioral claim, and if the claim is never examined, the model's conclusions about intervention effects rest on it silently. The third is design by platform: letting the conveniences of a particular tool shape the model. NetLogo's grid of patches makes it natural to represent space as a square lattice with neighbors defined by adjacency, which is rarely how human contact works. Python libraries that make networks easy to generate tempt modelers to use a standard random graph whether or not it resembles the population. The platform should be chosen after the design, to suit it, rather than the reverse. The fourth is outcome shopping: adjusting the model until it produces a result the modeler expected or wanted, and then reporting that version. Every modeler iterates, and not every iteration needs reporting, but when a structural choice changes the conclusion, the reader deserves to know that the conclusion depends on it. Pre-specifying the question, the target patterns, and the main analyses before extensive exploration, much as a trial protocol pre-specifies outcomes, is a discipline that agent-based modeling has been slow to adopt and would benefit from. Design as an iterative discipline None of this implies that design is a single phase completed before building begins. In practice, the question, the scope, the ODD description, and the code evolve together. A first prototype reveals that some mechanism matters more than expected, or that a target pattern cannot be reproduced without an additional process, or that a parameter the modeler thought unimportant drives everything. Each discovery sends the modeler back to the design. What matters is that each iteration is anchored to the purpose. A model that grows because the modeler found something interesting to add, rather than because the question required it, becomes harder to understand with every addition. A model that grows because a target pattern demanded a new mechanism, and that keeps its ODD description current as it does, grows in a way that a reader can follow and a critic can challenge. That is the standard every model in social epidemiology should meet: not that it is realistic, but that its reasons for being as it is are stated and defensible. Chapter 3. Who Meets Whom: Contact Networks, Places, and Synthetic Populations A pathogen cannot travel between two people who never come close to each other, and a behavior rarely spreads between two people who never communicate. The structure of contact is therefore the skeleton on which every agent-based model of contagion hangs. Get it roughly right and much else can be approximate; get it badly wrong and no refinement of the disease or decision model will rescue the results. This chapter examines how contact structure is represented, what properties of it matter most, and how large models build populations that resemble real ones. What we know about human contact For most of the history of infectious disease modeling, contact patterns were inferred rather than measured. Modelers assumed mixing matrices, often with arbitrary parameters for how much people mixed within versus between age groups, and tuned them until results looked plausible. That changed with a series of empirical studies, of which the most influential was POLYMOD. The POLYMOD contact survey, reported by Joël Mossong and colleagues in PLoS Medicine in 2008, asked 7,290 participants in eight European countries to record every person they had contact with over a single day, defined as either a two-way conversation of at least three words in the other's physical presence or skin-to-skin contact. The participants recorded 97,904 contacts. The findings reshaped the field. Contact patterns were strongly assortative by age, with people mixing most with others their own age; school-age children and adolescents had the highest numbers of contacts; and there was a clear secondary band of contact between children and adults of parental age, reflecting households. Contacts that were physical and long in duration tended to occur at home, while contacts at work, school, and leisure were more numerous but more varied. Simulations using the measured patterns predicted that school-age children and young adults would experience the highest incidence early in the spread of a new respiratory pathogen, a prediction consistent with what was observed in the 2009 influenza pandemic. Because contact surveys are expensive, Kiesha Prem, Alex Cook, and Mark Jit developed synthetic contact matrices in 2017, using the POLYMOD data together with demographic and household data to project contact patterns by age and setting (home, school, work, and other) for 152 countries, many of which had never conducted a contact survey. The matrices were later updated, and they became standard inputs for COVID-19 models in low- and middle-income countries. Contact surveys were also repeated during the pandemic itself; the CoMix study in the United Kingdom and elsewhere documented how contacts fell dramatically during lockdowns and recovered unevenly afterward. Other studies measured contact directly with wearable sensors. Salathé and colleagues, in a 2010 paper in the Proceedings of the National Academy of Sciences, equipped nearly everyone in an American high school with wireless sensors for a day and recorded every close-proximity interaction. The resulting network was dense and highly structured, with most contact concentrated in a relatively small number of repeated, long-duration interactions; simulations on the measured network predicted outbreak dynamics that differed from those on randomized versions of the same network. The SocioPatterns collaboration has made similar measurements in schools, hospitals, and workplaces, and made much of the data publicly available. Several lessons from this body of work matter for model design. Contact is not uniform: it varies by age, setting, day of the week, season, and circumstance. Contact is repeated: most of a person's contacts on a given day are people they also saw yesterday. Contact is clustered: the people you meet tend to meet each other. And contact is heterogeneous in intensity: a few hours together at home is a different exposure from a brief exchange at a shop. A model that treats contact as random encounters of equal intensity misses all four. The network properties that matter When contact is represented as a network, with agents as nodes and relationships as edges, a few structural properties dominate epidemic dynamics. The comparison of standard network models in Table 2 below is organized around them. Degree distribution describes how many contacts each agent has. What matters for transmission is not only the mean but the spread. A fundamental result in network epidemiology is that, for an infection spreading on a static network, the relevant measure of how many new infections an infected person produces depends on the ratio of the mean squared degree to the mean degree, not on the mean degree alone. Intuitively, people with many contacts are both more likely to become infected and more likely to pass infection on, so heterogeneity in degree amplifies transmission. The friendship paradox, described by the sociologist Scott Feld in 1991, captures a related point: on average, your friends have more friends than you do, because highly connected people appear in many people's friendship lists. Infection reaches highly connected people early, which is why early growth can be faster than population averages suggest and why it can slow once the most connected have been infected or immunized. In 1999 Albert-László Barabási and Réka Albert showed, in Science, that a simple process of growth with preferential attachment produces networks whose degree distributions follow a power law, so-called scale-free networks. In 2001 Romualdo Pastor-Satorras and Alessandro Vespignani showed in Physical Review Letters that on infinite scale-free networks the epidemic threshold vanishes: any positive transmission probability allows an infection to persist. The result was important, but its practical reach has been debated. Claims that particular human contact networks, such as sexual networks, are truly scale-free were contested on statistical grounds, and real networks are finite, which restores a threshold. The more robust lesson is that high variance in degree, whatever its exact distribution, lowers the threshold for spread and concentrates transmission on a minority of highly connected people. Clustering measures how often two contacts of the same person are also in contact with each other. Clustering tends to slow epidemic spread for simple contagions, because infections within a tight cluster often reach people who have already been infected by someone else, wasting transmission opportunities. Duncan Watts and Steven Strogatz showed in Nature in 1998 that a lattice with a small number of random long-range links can combine high clustering with short average path lengths, the small-world property. For contagion, the long-range links matter disproportionately: they carry infection between otherwise separate clusters. Assortativity measures whether similar nodes connect to each other. Age assortativity is strong in human contact, as POLYMOD showed. Socioeconomic and ethnic assortativity are also strong, reflecting what Miller McPherson, Lynn Smith-Lovin, and James Cook, in a much-cited 2001 review in the Annual Review of Sociology, called homophily: the tendency of similar people to associate. For social epidemiology, assortativity is where contact structure meets inequality. If people at high risk of exposure mostly contact others at high risk, infection concentrates in their communities; if they contact widely, it spreads out. Temporal dynamics describe how the network changes. Sexual partnerships form and dissolve; workplaces empty on weekends; schools close for holidays; lockdowns delete entire layers of contact. Martina Morris and Mirjam Kretzschmar showed in the 1990s that concurrent sexual partnerships, overlapping in time, could greatly accelerate HIV spread compared with the same number of partnerships occurring serially, because concurrency links people into larger connected components at any moment. Static networks miss this entirely. Table 2. Common network representations and what they preserve. Network type How it is generated Degree variance Clustering Typical use Random (Erdős–Rényi) Each pair linked with fixed probability Low Very low Baseline comparison only Small-world (Watts–Strogatz) Lattice with a fraction of links rewired Low High Local clustering with long-range links Scale-free (Barabási–Albert) Growth with preferential attachment Very high Low Exploring hub effects Configuration model Random wiring to a specified degree sequence As specified Low Matching observed degree data Exponential random graph Statistical model fitted to network data As fitted As fitted Empirically grounded sexual and social networks Multilayer settings Households, schools, workplaces, community Emerges from settings High within settings Large respiratory-disease models Places as generators of contact An alternative to specifying a network directly is to specify places, and let contact arise from shared presence in them. Agents belong to a household, perhaps a school or workplace, and visit community locations such as shops, transit, and places of worship. Contact occurs among people in the same place at the same time. This is how most large respiratory-disease models represent contact, and it has several advantages. It matches how data are collected. Censuses describe household composition; education statistics describe school enrollment; labor surveys describe employment by industry and firm size. It matches how interventions work: closing schools, restricting workplaces, and limiting gathering sizes act on places, not on abstract network edges. And it naturally produces realistic clustering, because everyone in a household is in contact with everyone else, and classmates are largely in contact with each other. Covasim, the model developed by Cliff Kerr and colleagues at the Institute for Disease Modeling and described in PLoS Computational Biology in 2021, organizes contact into layers corresponding to households, schools, workplaces, and the community, each with its own number of contacts and transmission weight. An intervention such as school closure removes or reduces the school layer; working from home reduces the workplace layer for some agents. The same layered structure appears in OpenABM-Covid19, developed by Robert Hinch and colleagues at Oxford, and, in more elaborate form, in national-scale models such as the Imperial College model and FRED (the Framework for Reconstructing Epidemic Dynamics) developed at the University of Pittsburgh. The place-based approach has its own weaknesses. Not everyone in a workplace contacts everyone else, so large workplaces must be subdivided or given within-place contact networks. The community layer, which captures all contact outside the structured settings, is usually the least well measured and often ends up as a random-mixing catch-all whose parameters are tuned to fit data. And the approach tends to omit the settings where social disadvantage concentrates exposure: multigenerational and overcrowded housing, shift work in facilities such as meatpacking plants and care homes, crowded public transit, prisons, and shelters. These are exactly the settings social epidemiology most needs to represent. A model that treats all workplaces as statistically identical, differing only in size, cannot show why a warehouse worker faced a different risk from an office worker who could move online. The remedy is to build occupational and residential structure into the population itself. Agents can be assigned occupations from labor data, with each occupation carrying attributes such as the possibility of remote work, typical workplace density, and whether the job involves contact with the public. Households can be assigned from census data preserving crowding, generational composition, and income. Contact within these settings can then be scaled by the relevant attributes. The data exist for many high-income countries; the modeling choice to use them is what determines whether the resulting simulation can say anything about inequality. Synthetic populations Large agent-based models require a population of agents whose joint distribution of attributes resembles a real population. Real microdata at the individual level are rarely available for whole populations, because of privacy constraints, so modelers construct synthetic populations: artificial people whose aggregate characteristics match published tables. The standard technique descends from work by Richard Beckman, Keith Baggerly, and Michael McKay, published in Transportation Research Part A in 1996 for transportation planning. It combines two data sources: a sample of complete individual and household records, such as the Public Use Microdata Sample from the American Community Survey, which preserves the joint distribution of attributes but only for a small fraction of the population; and aggregate tables for small areas, such as census tract counts by age, household size, and income, which cover the whole population but only as marginal totals. Iterative proportional fitting adjusts weights on the sample records until their weighted totals match the small-area marginals, and households are then drawn from the weighted sample to populate each area. The result is a population in which every tract has approximately the right number of households of the right sizes and income levels, and in which each household has plausible internal structure drawn from real records. Several refinements are now common. Combinatorial optimization methods select households to match multiple constraints simultaneously. Methods based on Bayesian networks or deep generative models learn the joint distribution of attributes and sample from it, which can produce combinations absent from the sample. Once people and households exist, they must be assigned to schools and workplaces, typically using enrollment and commuting data; the SynthPops library, developed alongside Covasim, does this for school, workplace, and long-term-care-facility networks. The RTI International synthetic population for the United States, built for the National Institutes of Health's MIDAS modeling network, was one of the earliest widely shared examples. Synthetic populations raise their own methodological concerns. They reproduce the tables they were fitted to, but not necessarily the correlations those tables do not contain; if the target data do not cross-classify income with occupation, the synthetic population may get their joint distribution wrong. They are snapshots, while real populations move, age, and change households. And they inherit the biases of their source data, including undercounts of precisely the populations social epidemiology cares about: people who are homeless, undocumented, incarcerated, or living in informal housing. A modeler should know which attributes in a synthetic population were fitted, which were imputed, and which were simply assumed. Mobility, space, and the geography of exposure Contact also depends on where people go. For much of the history of spatial epidemic modeling, movement between places was represented with gravity models, in which the flow between two locations rises with their populations and falls with the distance between them, or with commuting matrices from censuses. Since the late 2000s, mobile phone records and location data from smartphone applications have made it possible to observe movement far more directly, and during the COVID-19 pandemic such data became central to many models. The most instructive example for social epidemiology is the study by Serina Chang, Emma Pierson, Jure Leskovec, and colleagues published in Nature in 2021. Using anonymized, aggregated location data, they built hourly networks linking census block groups to specific points of interest, such as restaurants, grocery stores, gyms, and places of worship, across ten of the largest metropolitan areas in the United States in the spring of 2020. A relatively simple infection model, run on these mobility networks, reproduced observed case trajectories, and it attributed a large share of infections to a small minority of points of interest, typically crowded venues where people stayed a long time. It also generated socioeconomic and racial disparities in infection that were not built into the model's transmission rules. The model predicted higher infection rates among lower-income and non-white neighborhoods for two reasons that the mobility data revealed: residents of those neighborhoods had reduced their mobility less, consistent with jobs that could not be done from home, and the places they visited, such as grocery stores, were on average more crowded and associated with higher risk than the same categories of place visited by residents of wealthier neighborhoods. Strictly speaking, that model was a metapopulation model on a bipartite network rather than a full agent-based model, since it tracked neighborhoods rather than individuals. But its lesson transfers directly. Inequality in exposure arose from structure, from where people had to go and how crowded those places were, rather than from any difference in how the pathogen treated different groups. A model that assigned everyone the same contact pattern, however detailed, would have missed the disparities entirely. A model that represented the structure produced them without being told to. Mobility data carry their own cautions. Samples of smartphone users are not representative, and they undercount older people, people without smartphones, and people who disable location services, groups that are themselves unevenly distributed across social strata. Commercial data sources change their methods and availability without notice, which threatens reproducibility; several of the datasets widely used in 2020 were later restricted or discontinued. And location proximity is an imperfect proxy for transmission-relevant contact, since two people in the same store for twenty minutes may or may not have come close enough to matter. Used with those limits in mind, movement data are among the most powerful tools available for grounding the spatial structure of a model in observation rather than assumption. Choosing a representation How detailed should contact structure be? The pattern-oriented principle from Chapter 2 gives the answer: detailed enough to reproduce the patterns the question depends on. Several guidelines follow. For questions about the overall size and timing of a respiratory epidemic in a general population, an age-structured contact matrix, or a simple household-plus-community structure, often suffices, and more detail changes little. For questions about household-level interventions, such as isolation of cases within the home or household quarantine, households must be explicit. For questions about school or workplace policy, those settings must be explicit and appropriately sized. For sexually transmitted infections and blood-borne infections, where contact is sparse, partnership duration and concurrency matter enormously, and a dynamic network model is usually required; the EpiModel package for R, described by Samuel Jenness, Steven Goodreau, and Martina Morris in the Journal of Statistical Software in 2018, fits temporal exponential random graph models to partnership data for this purpose. For questions about inequality, the settings and attributes that generate differential exposure must be present, which usually means occupation, housing density, and residential segregation. Whatever representation is chosen, it should be tested against alternatives. A useful exercise is to rerun the model with the network randomized while preserving each agent's number of contacts, which removes clustering and assortativity but keeps degree. If the results change substantially, structure beyond degree matters, and the representation of that structure deserves close scrutiny. If they do not, a simpler representation may be adequate. Either way, the modeler learns something about the question that no single configuration could reveal. Hashtags: #AgentBasedModelingInSocialEpidemiology #AgentBasedModeling #SocialEpidemiology #ContagionModeling #HumanBehavior #InfectiousDiseaseModeling #BehavioralContagion #ComplexSystems #SEIRModel #ContactNetworks #NetworkEpidemiology #SyntheticPopulations #ODDProtocol #AdaptiveBehavior #StochasticSimulation #SocialDeterminantsOfHealth #HealthInequality #SpatialEpidemiology #MobilityModeling #VaccineBehavior #ComplexContagion #Homophily #SensitivityAnalysis #PolicySimulation #FutureOfComputationalEpidemiology
Latest Book Releases:










































