top of page

Welcome to the VBNN Digital Library

Unlock a Vast Knowledge Ecosystem

Featuring over 30,000 books, academic papers, illustrations, and expert insights—continuously updated to support your research and professional growth.

​

Welcome to our library!

Here, you will find an exclusive collection created 100% by our own faculty, meaning you will not find these resources anywhere else. Over the last 20 years, our team has written much more than what is currently online, and we are actively working to upload our complete back catalog. We update our platform regularly, so be sure to check back from time to time. If you ever need help finding a specific resource, you can always contact us!

​

Maximize Your Access

Log in to instantly view and download tailored resources directly aligned with your specific program and curriculum.

Ready to begin? Sign in above to explore your personalized dashboard.

​

Please note: Login is only possible using your institutional email address; otherwise, the system will not recognize your account.

​

VBNN Library AI

Introducing our fully integrated Library AI. Designed to support your research, you may submit inquiries in any language and receive precise, evidence-based responses drawn exclusively from our published scholarly articles and textbooks.

Search...

Latest Publications:

Search this site

Results found for empty search

  • Secondary Data Analysis - Mining Public Registries, Censuses, and Longitudinal Biobanks)

    Download the Book (PDF): Introduction Somewhere in a mobile examination trailer parked outside a county fairground, a technician draws blood from a volunteer, labels the tube, and logs the time. Somewhere in a Swedish tax office, a clerk updates an address. In a warehouse in Stockport, a robot retrieves one of several million frozen aliquots so that a machine in another building can count five thousand proteins in it. In 1950, an enumerator in rural Georgia wrote down the occupation of a man he would never see again, and in April 2022, seventy-two years later, that page became a row in a public file. None of these people was working for you. Every one of them made decisions that will shape the numbers you report. That is the central fact of secondary data analysis, and it is easy to forget once the files are on your screen. The UK Biobank, the National Health and Nutrition Examination Survey (NHANES), the Integrated Public Use Microdata Series (IPUMS), the Nordic population registries, and their many cousins are extraordinary public goods. They let a doctoral student test a hypothesis on half a million people for a few thousand pounds, let a historian follow individuals across four decades of censuses, and let a nutrition scientist estimate the prevalence of iron deficiency in American toddlers with a precision that no single laboratory could ever afford. They have also produced a large and growing literature of confident, precisely estimated, wrong answers. The two facts have the same cause. These datasets are cheap to analyse because someone else already bore the cost of designing and collecting them, and they are easy to misanalyse because the analyst did not make, and often does not understand, the design decisions baked into every row. The controlling idea This book argues one thing: a secondary dataset is the output of a data-generating process that someone else designed, and valid analysis begins by reconstructing that process and building your analysis to respect it. Sampling probabilities, oversampling of particular groups, the reasons people agreed or refused to take part, the way records were stitched together across institutions, the legal terms under which the data were released, and the variables that were never collected at all are not background details. They are part of the data. An analysis that ignores them is not a simplification of the correct analysis. It is a different analysis, answering a question nobody asked about a population that does not exist. Put that way, the idea sounds obvious. In practice it is violated constantly, for understandable reasons. The files arrive looking like any other rectangular dataset. Standard statistical software will happily fit an unweighted logistic regression to NHANES and return a tidy table of odds ratios and standard errors, without any warning that the standard errors are too small, the intercept describes nobody, and the sample over-represents the groups the survey designers deliberately oversampled. A Cox model fitted to the UK Biobank will return hazard ratios with confidence intervals so narrow they look like physical constants, and nothing in the output will mention that only about one in eighteen invited people chose to take part, or that the ones who did smoke less, drink less, and die less than the population they came from. Software reports what you asked it to compute. It does not report what you should have asked. Three families of data The book deals with three broad families of secondary data, and it is worth naming them at the start, because each carries a characteristic design problem. The first is the probability survey, of which NHANES is the leading health example and the American Community Survey and the Current Population Survey (both accessible through IPUMS) are the leading social examples. These are built from explicit sampling designs, with stratification, clustering and unequal selection probabilities, and they come with weights and design variables that encode that design. Their characteristic problem is that the analyst must carry the design into the analysis, and most introductory training does not teach how. The second is the register or census: the complete, or nearly complete, enumeration of a population for administrative or constitutional purposes. The Danish Civil Registration System, which has assigned a unique personal identification number to every resident since 1968, and the United States decennial censuses, now released in full-count form through 1950, are archetypes. Their characteristic problem is not sampling but meaning. The variables were recorded to administer benefits, collect taxes or apportion seats in a legislature, and they measure what those purposes needed, not what a researcher needs. Linking them across time and across institutions introduces a second problem: every link is a claim that two records belong to the same person, and some of those claims are false. The third is the volunteer cohort or biobank. The UK Biobank recruited around 500,000 people aged 40 to 69 between 2006 and 2010, measured them extensively, stored their biological samples, and has since linked them to hospital, cancer, death and, for a large subset, primary care records, while adding genotyping, whole-genome sequencing, imaging on a subsample and, from 2026, large-scale proteomics. Its characteristic problem is selection. Participants were invited, not sampled with known probabilities, and the few who accepted differ systematically from those who did not. That matters a great deal for some questions and very little for others, and telling which is which is one of the more demanding skills in modern epidemiology. Cutting across all three families is the problem that no amount of sample size can fix: confounding by variables the dataset never measured. Secondary data are, by definition, collected for someone else's purposes. The confounder you most need is often the one nobody thought to record. How the book is organised Chapter 1 establishes the habit of mind the rest of the book depends on: reading a dataset as the product of decisions, and reading its documentation before its data. It also compares the major resources side by side. Chapter 2 deals with access and governance, which is not paperwork to be delegated but a set of constraints that determine what analyses are possible, where they can be run, and what can be published. Chapters 3 and 4 cover complex survey data: first the mechanics of weights, strata, clusters and variance estimation, using NHANES as the running example, and then the harder questions of when and how to weight in regression models, with missing data and in pooled analyses. Chapter 5 turns to volunteer cohorts and the problem of selection, including the recent work on reweighting the UK Biobank. Chapter 6 covers record linkage, from deterministic rules through probabilistic matching to the historical census linking that has transformed economic history, and the biases that linkage error introduces. Chapters 7 and 8 address unmeasured confounding: first the study designs that can blunt it, and then the quantitative sensitivity analyses that tell you how much an unmeasured confounder would have to do to overturn your result. The Conclusion draws out what follows for practice. Who this is for, and what it assumes The intended reader is someone who has fitted a regression model and read an epidemiology or social science paper critically, and who now wants to work with large public datasets without falling into the traps that catch capable people. That includes graduate students, clinicians moving into research, analysts in government and industry, and senior researchers trained before these resources existed. The book assumes familiarity with ordinary regression and basic probability, but no survey sampling theory, no causal inference formalism, and no particular software. Where software matters, I name the relevant tools: the survey package in R, the svy prefix in Stata, PROC SURVEYREG and its relatives in SAS, and the cloud platforms on which some of these datasets now must be analysed. A note on currency. The access arrangements for these datasets change often, and several changed materially between 2023 and 2026. The UK Biobank no longer routinely allows bulk downloads; analysis happens on its cloud Research Analysis Platform. NHANES released its first post-pandemic cycle, covering August 2021 to August 2023, with a redesigned sample and much lower response rates, and the programme itself passed through a period of acute institutional uncertainty in late 2025. Europe adopted a regulation creating a European Health Data Space whose secondary-use provisions will begin to apply in 2029. The details given here were checked against the data custodians' own documentation in 2026. They will drift, and the book tries to teach the questions to ask rather than only the answers that are true today. A word on what "public" means The word "public" in public data is doing a lot of work, and it means different things for different resources. NHANES public-use files can be downloaded by anyone, anonymously, without an application. IPUMS extracts require free registration and agreement to terms of use, and the most detailed international and full-count files carry additional restrictions. The UK Biobank is open to any bona fide researcher for health-related research in the public interest, but requires an application, a fee, and analysis on a controlled platform. The Nordic registries are open, in the sense that researchers from anywhere can in principle obtain them, but only through national authorities, often only for analysis on secure servers inside the country, and usually only after months of review. These differences are not incidental. They reflect different judgments about the balance between the value of research and the risk to the people whose lives the data describe, and an analyst who understands those judgments will design better studies and have fewer of them rejected. The people in these datasets did not consent to your specific study. In most cases they consented, or were compelled by law, to a broad purpose, and trusted the custodian to police the boundaries. The obligations that follow from that trust are part of the method, not an ethical appendix to it. A good secondary analysis is one that the participants, if they could read it, would recognise as a fair use of what they gave. Chapter 1. Reading the Data Before the Data The most consequential hour in a secondary analysis is usually spent before any data are loaded. It is spent reading: the design documentation, the codebook, the analytic guidelines, the release notes, and the list of known errors. Analysts who skip this hour tend to spend the following six months discovering, one referee report at a time, what the documentation would have told them. This chapter is about how to read a secondary dataset as an artefact with a history, and about the handful of questions whose answers determine almost everything that follows. The data-generating process as the unit of understanding Every dataset is the endpoint of a chain of events. For a survey, the chain runs roughly as follows: a target population is defined; a sampling frame is chosen that approximates it; the frame is divided into strata and clusters; units are selected with known but unequal probabilities; some selected units cannot be found, some refuse, some complete only part of the protocol; answers and measurements are recorded, cleaned, edited, top-coded, imputed and suppressed; weights are computed to reverse the selection and adjust for nonresponse; and a public file is released with some identifying detail removed. For a registry, the chain runs through the administrative purpose that created the record, the rules that determined who must be recorded, the coding systems in force at each date, the incentives of the people doing the recording, and the processes by which records were corrected or linked. For a biobank, it runs through the invitation strategy, the decision to volunteer, the assessment protocol, the choice of which samples to assay and in what order, and the linkage of external records whose own chains are just as long. Each link in the chain can distort the relationship between what is in the file and what is true of the population. The analyst's first job is to know which links matter for the question at hand. A useful discipline is to write down, before looking at any results, the path by which a single individual in the target population ends up as a row in your analysis dataset, with a value for every variable you intend to use. Every step on that path at which the probability of reaching the next step could depend on the exposure, the outcome, or both, is a place where bias can enter. Some of those steps will be handled by the survey weights. Some will be handled by design choices you make. Some will be handled by nothing at all, and those are the ones to state honestly as limitations. Consider a concrete case. Suppose you want to estimate the association between serum vitamin D concentration and depressive symptoms among American adults, using NHANES. A person enters your analysis only if their county was selected as a primary sampling unit, their household was selected within the county, they were selected within the household, they completed the home interview, they attended the mobile examination centre, they consented to and completed the blood draw, their sample was successfully assayed for 25-hydroxyvitamin D, and they completed the Patient Health Questionnaire (PHQ-9) module in the private interview. The first three steps are controlled by the design, with known probabilities that the weights reverse. The fourth and fifth are affected by nonresponse, which the weights attempt to adjust for using demographic variables. The sixth, seventh and eighth are item-level losses that the standard weights may not address at all. If people with depression are less likely to attend the examination centre, or less likely to complete the blood draw, the weights will not know. That does not make the study impossible. It makes it a study whose limitations can be precisely described. What to read, and in what order The documentation for the major resources is extensive and uneven, and the order in which you read it matters. A sensible sequence is: the design overview, the analytic guidelines, the documentation for each specific file you will use, the release notes and errata, and finally a selection of well-cited papers that have used the same variables. For NHANES, the design overview and analytic guidelines are published for each cycle by the National Center for Health Statistics (NCHS). The data files are organised by component (demographics, dietary, examination, laboratory, questionnaire) and by cycle, with a letter suffix marking the cycle: files from 2017 to 2018 carry the suffix J, the combined pre-pandemic files covering 2017 to March 2020 carry the prefix P, and the August 2021 to August 2023 files carry the suffix L. Each file has its own documentation page listing the eligible sample, the protocol, the data processing and editing, and, crucially, which weight to use. That last point is not a formality. The laboratory components are often measured on a random subsample of examinees, with a subsample weight that must be used in place of the examination weight, and a fasting subsample carries its own weight again. For the UK Biobank, the central reference is the Data Showcase, which catalogues every data field by a numeric field identifier, with the number of participants who have data, the instance (the assessment visit, where 0 is baseline, 1 is the first repeat assessment in 2012 to 2013, 2 is the imaging visit that began in 2014, and 3 is the repeat imaging visit that began in 2019), and the array index where a question allowed multiple answers. The Showcase also hosts resource documents describing each assay, each linked dataset and each derived outcome. Many analysts go straight to the phenotype they want and never read the resource document explaining how it was derived. That is how a study ends up treating the absence of a hospital diagnosis code as evidence that a condition is absent, in participants whose primary care records were never linked. For IPUMS, the key documents are the variable descriptions and the comparability discussions attached to each harmonised variable. IPUMS exists precisely to harmonise: to recode occupation, industry, education, relationship to household head and hundreds of other variables so that they mean approximately the same thing across censuses a century apart or across countries with different statistical traditions. The word "approximately" is doing work. The comparability notes explain where the harmonisation is exact and where it is a best effort, and they record changes in question wording, universe (who was asked the question) and coding that can create artificial trends. The universe statement for each variable is particularly important: many census questions were asked only of people above a certain age, or only of household heads, or only in a sample of households, and a variable coded as "not in universe" is not the same as a zero. For national registries, the key documents are usually validation studies published by the custodians or by academic users. The Danish registries, for example, have an unusually well-developed literature assessing the positive predictive value of specific diagnosis codes against medical record review. Reading those studies tells you which codes can be trusted to identify a condition and which cannot, and whether a code's meaning changed when the country moved from the eighth to the tenth revision of the International Classification of Diseases in 1994. Comparing the major resources The resources this book returns to most often differ along several dimensions at once: how participants entered the data, how many there are, who can get access and how, and what kind of error most threatens a naive analysis. Table 1 sets out those differences for five representative resources. The characteristic hazard in the final column is not the only problem each resource has, but it is the one that most often invalidates published work. Table 1. Five major secondary data resources compared (sources: custodians' documentation, 2025–2026). Resource How people entered Scale Access route Characteristic hazard NHANES (US) Multistage probability sample, two-year cycles About 5,000 examined per year pre-2020; 8,860 examined in 2021–2023 Public-use download; restricted files via NCHS Research Data Center Ignoring weights, strata and clusters IPUMS USA and CPS Census and survey samples, plus full-count censuses 1850–1950 Samples of 1–5% of population; full counts of tens of millions Free registration and terms of use Harmonisation artefacts; "not in universe" coding UK Biobank Invited volunteers aged 40–69, recruited 2006–2010 About 500,000 Application, tiered fee, cloud platform analysis Healthy volunteer selection; linkage completeness Nordic registries Legal obligation to register all residents Whole national populations National authority approval; remote secure servers Administrative coding changes; unmeasured lifestyle confounders Linked administrative data (e.g., hospital episodes) Contact with a service Everyone who used the service Custodian approval, often via trusted research environments Selection into care; linkage error Two features of the table deserve comment. First, scale and representativeness run in opposite directions for the health resources. The survey with the explicit probability design, NHANES, is the smallest. The biobank with the largest sample, the UK Biobank, has the least representative recruitment. The registries escape the trade-off by covering everyone, but pay for it in the thinness of what they record: a Danish registry knows your prescriptions and hospital diagnoses in detail but does not know whether you smoke. Second, the access routes are becoming more restrictive over time at the top end and more open at the bottom. Public-use files remain freely downloadable, but the richest data increasingly live in trusted research environments where the analyst brings code to the data rather than data to the code. Chapter 2 examines what that shift means in practice. Matching the question to the data Before choosing variables and methods, it is worth being explicit about what kind of question you are asking, because the three main kinds of question make different demands on the data-generating process. A descriptive question asks what is true of a defined population: the prevalence of diagnosed and undiagnosed diabetes among American adults in 2021 to 2023, the share of married women in the labour force in 1950, the age-standardised incidence of hip fracture in Denmark. Descriptive questions demand representativeness. The estimate is only meaningful for the population the sample represents, and the whole apparatus of survey weights exists to make that link. For descriptive questions, a volunteer cohort such as the UK Biobank is a poor instrument, and its own investigators have always said so. The prevalence of smoking in the UK Biobank tells you about UK Biobank participants. A predictive question asks how well some information forecasts some outcome: whether a polygenic score improves the prediction of coronary disease beyond conventional risk factors, whether a set of blood proteins predicts ten-year mortality. Prediction does not require causal interpretation of the coefficients, but it does require that the relationship between predictors and outcome in the development data resembles the relationship in the population where the model will be used. Selection matters here in a subtler way. A risk model developed in a population healthier than the target will usually be miscalibrated, predicting too few events, even if its discrimination is fine. A causal question asks what would happen to an outcome if an exposure were changed: whether statins reduce dementia risk, whether a minimum-wage rise reduces employment, whether moving out of a high-poverty neighbourhood improves children's adult earnings. Causal questions demand that the comparison between exposed and unexposed be free of confounding, selection bias and measurement error sufficient to overturn the conclusion. Representativeness is not required for internal validity, though it matters for transporting the result to other populations. Chapters 5, 7 and 8 are mostly about causal questions, because that is where secondary data are most often used and most often misused. Many published papers slide between these three kinds of question without noticing. A paper reports an association in the UK Biobank, interprets it causally in the discussion, and then recommends a population-wide policy on the strength of it, implicitly claiming descriptive validity for the whole country. Each of those three moves requires a separate justification. Naming which question you are asking, in a single sentence at the top of your analysis plan, prevents a surprising amount of confusion later. A pre-analysis protocol With the documentation read and the question named, the final step before analysis is to write a short protocol. For secondary data, the protocol has an additional job beyond its usual role in preventing selective reporting: it forces you to make design decisions explicitly, before you see which ones produce the most attractive results. The protocol for a secondary analysis should specify at minimum the following. The target population and how it maps onto the dataset's sampled or recruited population. The eligibility criteria, applied in order, with the expected number excluded at each step, so that the flow of participants can be reported and checked. The exposure, outcome and covariate definitions, down to the specific field identifiers, variable names or diagnosis codes. The time structure: when eligibility is assessed, when exposure is measured, when follow-up begins and ends, and what happens at the boundaries. The weights, strata and cluster variables to be used, and the variance estimation method. The handling of missing data. The primary analysis and the sensitivity analyses, including those addressing unmeasured confounding. Several registries and repositories now accept protocols for observational studies. The Open Science Framework hosts many; the UK Biobank requires a description of the research as part of the access application, though that description is not a protocol in the full sense. Registration matters more for secondary data than for primary data, not less. When the data already exist, the only barrier between an analyst and a hundred different specifications is self-discipline. With 500,000 participants and thousands of variables, some specification will produce a small p-value for almost any hypothesis. The protocol is the evidence that you chose yours first. A final habit is worth building at this stage. Before running the analysis you care about, reproduce a published number from the documentation or from a well-known paper using the same data. NCHS publishes prevalence estimates with standard errors in its Data Briefs; if your code cannot reproduce the published prevalence of obesity for the cycle you are using, to the decimal place and with the same standard error, something is wrong with your weights, your design variables, or your understanding of the eligible population. IPUMS publishes population totals that a correctly weighted extract should match. For the UK Biobank, the cohort profile papers report baseline characteristics that a correctly specified extract should reproduce. The replication step takes an hour and catches errors that would otherwise surface only after publication, if at all. Data that move under your feet A last feature of secondary data that the documentation reveals, and that newcomers rarely anticipate, is that the data change. A published NHANES file is occasionally re-released with corrections, and the documentation page records the date and nature of each revision. A laboratory variable may be recalibrated when an assay method changes, and NCHS may publish a crosswalk equation so that values from different cycles can be compared; an analyst who pools cycles without applying the crosswalk will find a spurious trend. IPUMS releases new versions of its databases, with new samples, corrected codes and revised harmonisation, and asks users to cite the specific version. An extract downloaded in one year may not match an extract of the same variables downloaded two years later. The UK Biobank changes in a further way. Participants may withdraw at any time, and when they do, the resource periodically notifies approved researchers, who are required to remove the withdrawn participants' data from their analyses. Hospital, death and cancer registry linkages are refreshed at intervals, extending follow-up and sometimes revising earlier records as late-registered events arrive. The censoring date for each linked source differs by country of the United Kingdom, because the records come from separate English, Scottish and Welsh systems with separate update schedules. A study that sets a single administrative censoring date for all participants, later than the date to which one nation's records are complete, will undercount events in that nation and bias any comparison that correlates with where people live. The practical consequences are simple to state. Record the date and version of every extract. Keep the raw extract untouched and do all derivation in scripts, so that a refreshed extract can be pushed through the same pipeline. Check the release notes before each new analysis. And when you report results, state which release, which version and which follow-up end dates you used, so that someone trying to reproduce your numbers two years later knows what to ask for. These are not bureaucratic niceties. In a field where the same dataset is analysed by thousands of groups, they are what make disagreements between groups resolvable, rather than a matter of whose extract happened to be older. Chapter 2. Access, Agreements and the Rules of Use Researchers new to secondary data often treat access as an administrative hurdle to be cleared as quickly as possible so that the real work can begin. That view is expensive. The terms under which a dataset is released determine which variables you can see, at what level of geographic and temporal detail, whether you can link them to anything else, where your computation must run, what you are allowed to publish, and how long you may keep what you derived. Those are design constraints. A study planned without them in view will often be redesigned, sometimes after months of work, when the access terms turn out to forbid the linkage it depends on or the output rules suppress the subgroup it was built to study. Why the rules exist The people in secondary datasets are real, and most of them are identifiable in principle. That is true even of files stripped of names and addresses. A combination of a few ordinary attributes, such as exact date of birth, sex and a small geographic area, is unique for most of the population of a country, and sensitive attributes in the same row, such as a psychiatric diagnosis, an HIV test result or an income, then become attached to a person. Genomic data are identifying by nature. Custodians therefore face a trade-off between the scientific value of detail and the risk that detail enables re-identification, and different custodians strike it differently. The most widely used framework for thinking about that trade-off is the Five Safes, developed at the UK Office for National Statistics and set out by Tanvi Desai, Felix Ritchie and Richard Welpton. It asks whether the project is appropriate (safe projects), whether the researchers can be trusted to use the data properly (safe people), whether the environment prevents unauthorised use (safe settings), whether the data themselves carry a disclosure risk that is appropriate to the other safeguards (safe data), and whether the statistical results leaving the environment are non-disclosive (safe outputs). The value of the framework is that it treats the dimensions as substitutes. A dataset with high intrinsic disclosure risk can still be made available if the setting is locked down and the outputs are checked. A dataset released for anyone to download must be heavily de-identified, because none of the other safeguards applies. Almost every access regime a researcher will meet can be read as a particular setting of the five dials. NHANES public-use files turn "safe data" up to maximum, by removing geography below the national level, top-coding extreme values, and masking the design variables, and turn the other four down to nearly nothing: anyone may download them, anywhere, for any purpose consistent with a short use statement. The NCHS Research Data Center turns the dials the other way: the restricted variables, including county identifiers and exact dates, are available only to approved projects, analysed in a secure environment, with every output reviewed before release. The principal regimes in 2026 Table 2 summarises the arrangements for the resources this book discusses, as documented by the custodians in 2025 and 2026. The detail changes frequently, and the custodians' own pages are the authority, but the structure of each regime has been stable enough to plan around. Table 2. Access arrangements for major secondary data resources (sources: custodian websites, checked 2026). Resource Who may apply Cost to researcher Where analysis happens Notable conditions NHANES public-use Anyone None Anywhere No attempt to identify participants NCHS restricted data Approved projects Fees for Research Data Center use Secure physical or virtual enclave Output review before release IPUMS USA, CPS, International Registered users; International requires application None Anywhere No redistribution; citation required UK Biobank Bona fide researchers, health-related public-interest research Tiered access fee plus cloud compute and storage Research Analysis Platform (cloud) Downloads exceptional; return of derived data Nordic registries Researchers affiliated with approved institutions Agency processing fees National remote-access servers Ethics and data-protection approval; aggregate outputs only The UK Biobank deserves the fullest treatment, because its regime changed substantially in the early 2020s and the change affects how studies are designed. Access is open to researchers at academic, charitable, government and commercial institutions, for health-related research in the public interest, and applications are reviewed against that criterion rather than for scientific merit. The fee is set by data tier. According to the resource's published fee schedule in 2026, three years of access to Tier 1 (questionnaire, physical measures, health outcomes and linked records) costs £3,000; Tier 2, which adds biochemistry, genotypes, proteomics and other assays, costs £6,000; and Tier 3, which adds imaging, whole-genome and exome sequencing and other large-scale data, costs £9,000, with lower annual fees for extensions and an extra charge for each additional collaborating institution. Students and researchers in low- and middle-income countries pay a much reduced fee of £500 for three years. The larger change is where the analysis happens. Since the introduction of the UK Biobank Research Analysis Platform, built on the DNAnexus cloud, approved researchers analyse the data inside the platform rather than downloading it, and the resource considers applications to download data only in exceptional circumstances. The platform charges for computation, for data storage and for data egress (moving results out), with an initial credit for new users and, for a period, larger credits for new projects. The effect on study design is concrete. An analysis that would have been free on a university server now has a marginal cost for every hour of computation. Genome-wide analyses of whole-genome sequence data on hundreds of thousands of participants can cost thousands of pounds in compute alone if written carelessly and far less if written well. Budgeting for computation belongs in the grant application, and efficient code, such as running association tests on pre-filtered variant sets and using the platform's optimised tools rather than exporting data to generic software, has become a matter of money rather than taste. The UK Biobank also requires approved researchers to return the results of their research, including derived variables, to the resource, so that later researchers can use them. This makes the resource cumulatively richer, and it means that when you derive a complex phenotype, you are writing documentation for someone else, whether you intend to or not. NHANES remains the most open major health survey in the world, but its restricted data are more tightly controlled. Researchers who need geographic identifiers, exact examination dates, or linked data beyond what the public-use files include apply through the federal Standard Application Process, and analyse the data in an NCHS or Federal Statistical Research Data Center, physically or through a virtual desktop. The linked mortality files are an important middle case. NCHS links NHANES participants to the National Death Index and releases public-use versions, most recently with follow-up through 31 December 2019, with certain details perturbed to protect confidentiality; the restricted versions contain the unperturbed data. For many mortality analyses the public-use files are adequate, and NCHS has published comparisons showing close agreement with the restricted versions for common analyses. IPUMS is free, but not unconditional. Users register and agree to terms that, for IPUMS USA, prohibit redistribution of the data except for subsets needed to meet journal replication requirements, prohibit redistribution of the full-count census files entirely, and require citation of the data in a specified form that includes the version. IPUMS International, which distributes census microdata from national statistical offices around the world under agreements with each, adds an application and restricts use to the approved research and classroom purposes described in it. These are the terms on which national statistical offices were persuaded to release their data, and a breach by one user endangers access for everyone. The Nordic countries illustrate a model in which whole-population data are available in principle to researchers from anywhere, but only through a regulated process. In Denmark, researchers work on de-identified registry data through the remote-access servers of Statistics Denmark or the Danish Health Data Authority, with only aggregate results leaving the environment. In Finland, the Finnish Social and Health Data Permit Authority, Findata, created under a 2019 act on the secondary use of health and social data, issues permits for combining data from multiple custodians and provides a secure operating environment. Sweden requires ethical approval from the Swedish Ethical Review Authority before the registry holders, such as the National Board of Health and Welfare and Statistics Sweden, will link and release data. Processing times of six months to more than a year are common across the region and should be built into project timelines. The "safe people" dial has its own machinery. In the United Kingdom, access to de-identified government microdata under the research provisions of the Digital Economy Act 2017 requires researchers to be accredited by the UK Statistics Authority, which in practice means completing a short training course on disclosure control and the legal framework and passing an assessment. NCHS and the Federal Statistical Research Data Centers require background checks and sworn confidentiality obligations. These requirements feel like friction to an individual researcher, but they are what allow custodians to release more detailed data than they otherwise could, and a research group that keeps several accredited analysts on its staff can move much faster than one that must train someone for every new project. What a data use agreement actually says When access requires a formal agreement, the agreement usually contains a predictable set of clauses, and each one has design implications. A purpose limitation clause restricts use to the project described in the application. A change of research question may require an amendment. That is a reason to describe the project broadly enough to accommodate sensible sensitivity analyses and follow-on questions, but not so broadly that the reviewers cannot tell what you intend. A no re-identification and no linkage clause forbids attempts to identify individuals and forbids linking the data to other sources without permission. The second part catches many researchers by surprise. Merging a restricted dataset with area-level deprivation indices, air pollution estimates or hospital characteristics at a geographic level finer than the release allows can breach the agreement, even though no individual is identified, because it increases disclosure risk. A security clause specifies where data may be stored, who may see them, and how they must be protected. In trusted research environments, the environment enforces this. In agreements that allow data to be held locally, the clause usually requires encryption, restricted access and a named data custodian at the researcher's institution. An output control clause specifies the rules that results must meet before they leave a secure environment. The rules usually include minimum cell sizes, prohibitions on reporting minimum and maximum values, and restrictions on releasing residuals or other individual-level quantities. The minimum cell size varies: the US Centers for Medicare and Medicaid Services, for example, prohibits publishing any cell of fewer than eleven beneficiaries, and many European trusted research environments apply thresholds of five or ten with rounding. Output rules affect design most for rare exposures, rare outcomes and small subgroups. A study of a rare cancer in a specific age and ethnic group may be analysable inside the environment and yet impossible to report in the detail readers need. Checking the output rules before designing the tables avoids that outcome. A publication clause may require notification of publications, submission of manuscripts for review before submission, or specific acknowledgment text. Review clauses are usually for disclosure checking, not scientific censorship, but they add time. A retention and destruction clause specifies what happens to the data at the end of the agreement. Researchers who want to preserve their analytic datasets for replication must negotiate this explicitly, because the default is usually destruction. Ethics review and the law Secondary analysis of de-identified public data is often outside the scope of research ethics review, or exempt from it, but the determination belongs to the ethics committee or institutional review board, not to the researcher. In the United States, analyses of publicly available, de-identified data such as NHANES public-use files generally do not meet the regulatory definition of human subjects research under the revised Common Rule, and most institutions will issue a brief determination to that effect. Restricted data usually require review, and data covered by the HIPAA Privacy Rule may be released as a limited data set, with dates and some geographic detail retained, only under a data use agreement meeting the rule's requirements. In the United Kingdom and the European Union, the General Data Protection Regulation (and its UK counterpart) treats pseudonymised data as personal data, because the custodian can re-identify them. Research processing of health and genetic data requires both a lawful basis and a condition for processing special category data; for research, the relevant condition permits processing for scientific research purposes subject to safeguards, including data minimisation and, where possible, pseudonymisation. In practice, data custodians such as the UK Biobank, which relies on participants' broad consent together with these provisions, handle most of this analysis, but the researcher's institution remains responsible for its own processing and may need its own data protection impact assessment. The European Health Data Space regulation, which entered into force on 26 March 2025, will reshape access to health data across the European Union. Its secondary-use provisions, which create national health data access bodies empowered to issue data permits for research, policy and innovation and to make data available in secure processing environments, are scheduled to apply from March 2029, with some categories, including genetic data, following later. The model is close to the Finnish one: a single permit authority, secure processing environments, and outputs restricted to anonymous statistics. Researchers planning multinational European studies for the end of the decade should expect the permit process to become the standard route. Designing with the constraints The practical conclusion is that access terms should be read at the protocol stage, not after approval, and the analysis designed around them. If outputs must be aggregated, decide in advance which aggregates answer the question, and check that they will pass the cell-size rules. If the analysis must run in a cloud environment with metered computation, write and test the code on a small synthetic or public dataset first, and bring only debugged code into the paid environment. If linkage to external data is essential, name the external data in the application and obtain permission for the linkage explicitly. If the data will need to be retained for replication, negotiate retention at the start. It is also worth thinking about what a secure environment does to reproducibility. Code written inside a trusted research environment can usually be exported, but data cannot, and some environments restrict the software that can be installed. Openly publishing the code, together with the precise specification of the extract and the software versions, is the best available substitute for publishing the data. The OpenSAFELY platform, developed during the COVID-19 pandemic to analyse primary care records for tens of millions of patients in England, made that principle a requirement: all code executed against the real data is published, and every execution is logged publicly. It demonstrated that full transparency of method is compatible with complete confidentiality of data, and it is a model that other environments are gradually adopting. Finally, access regimes change, and the direction of change matters for how long a study takes. The broad trend since about 2015 has been away from releasing copies of sensitive data and towards bringing analysts into controlled environments. That trend is unlikely to reverse. It is safer for participants, it gives custodians visibility of how their data are used, and it lets them host datasets, such as whole-genome sequences on half a million people, that are too large to move. The cost falls on analysts, who must learn to work in environments that are less flexible than their own machines, and on the pace of research, which slows at every step of output review. Planning for that pace is part of the design. Chapter 3. Weights, Strata and Clusters: The Mechanics of Survey Data A national probability survey is a deliberately distorted picture of a population, together with the instructions for undoing the distortion. The distortion is intentional. Survey designers oversample small groups so that estimates for them are precise, cluster interviews geographically so that fieldwork is affordable, and stratify so that important subdivisions of the population are guaranteed representation. The instructions come in the form of weights, stratum identifiers and cluster identifiers. An analyst who uses the data without the instructions gets the distorted picture. This chapter explains what each part of the design does, how to carry it into an analysis, and where the common mistakes lie, using NHANES as the running example because its design is complex, well documented and representative of the major health examination surveys. Why the sample does not look like the population The NHANES sample is drawn in four stages: primary sampling units, which are counties or groups of contiguous counties; segments within counties, typically groups of census blocks; households within segments; and individuals within households. At each stage, units are selected with probabilities that the design controls, and those probabilities are deliberately unequal. For most of the continuous NHANES era, from 1999 until data collection was suspended in March 2020, the survey oversampled specific groups. In different periods these included Mexican American and later all Hispanic persons, non-Hispanic Black persons, non-Hispanic Asian persons from 2011, low-income non-Hispanic White and other persons, and older adults. Oversampling a group means selecting its members with higher probability than their share of the population. The result is that an unweighted NHANES sample contains far more Hispanic, Black and Asian participants, relative to their population share, than a simple random sample would. That is the point: it allows reliable estimates for those groups. It also means that any unweighted national estimate, such as the prevalence of hypertension among all American adults, is pulled toward the values in the oversampled groups. The August 2021 to August 2023 cycle changed the design in response to the pandemic. To reduce the amount of in-person contact needed during household screening, NCHS eliminated oversampling by race, Hispanic origin and income. Instead, within each selected household, all members aged 0 to 19 and all members aged 60 and over were selected, along with one or two randomly chosen adults aged 20 to 59, depending on household size. The sample remains unequal-probability, now by age and household composition rather than by race and income, and the weights still matter. The practical consequence NCHS draws in its analytic guidance is that estimates for some race and Hispanic origin subgroups in 2021 to 2023 are less precise than in earlier cycles, because those groups are no longer boosted in the sample. A small worked example shows how large the effect of oversampling can be. Suppose a population contains 90,000 people in group A and 10,000 in group B, and a condition affects 10 percent of group A and 30 percent of group B. The true population prevalence is (9,000 + 3,000) / 100,000, or 12 percent. A survey samples 900 people from group A (1 in 100) and 900 from group B (9 in 100). Unweighted, the sample contains 90 cases in A and 270 in B, for an apparent prevalence of 360 / 1,800, or 20 percent. Each respondent in A represents 100 people and each in B represents about 11.1. Weighting the cases by those factors gives 9,000 + 3,000 cases in a weighted total of 100,000 people, recovering the 12 percent. Real surveys have hundreds of distinct weight values rather than two, but the logic is identical. How weights are built A final survey weight is typically the product of three components, and knowing them helps in understanding what the weight can and cannot fix. The base weight is the inverse of the probability of selection. A person selected with probability 1 in 20,000 receives a base weight of 20,000, meaning that the person stands in for 20,000 people in the population. Base weights undo the deliberate distortions of the design, and nothing else. The nonresponse adjustment inflates the weights of respondents to compensate for selected people who did not take part. It is usually done within adjustment cells or by modelling response propensity from variables known for both respondents and nonrespondents, such as age, sex, race and Hispanic origin, and characteristics of the area. The adjustment corrects nonresponse bias only to the extent that, within cells defined by those variables, respondents resemble nonrespondents on the outcome being studied. If people with undiagnosed diabetes are less likely to attend the examination, and that tendency is not explained by the adjustment variables, the weights will not fix it. The calibration or post-stratification adjustment scales the weights so that weighted totals match independent population counts, usually from the Census Bureau's population estimates, for categories of age, sex, race and Hispanic origin. After calibration, a weighted NHANES sample reproduces the Census Bureau's count of, say, non-Hispanic Black women aged 40 to 59 exactly. NHANES provides separate weights for each stage of participation, because each stage has its own nonresponse. The interview weight applies to everyone who completed the household interview; the examination weight, which is always zero for interviewed-only participants, applies to everyone who also attended the mobile examination centre; and further weights apply to subsamples selected for particular laboratory tests, to the fasting subsample, and to participants who completed one or two days of dietary recall. For August 2021 to August 2023, NCHS also created phlebotomy weights for analyses of blood analytes, to account for examined participants who did not complete the blood draw. These last weights are an example of the survey adjusting for a stage of item-level loss that earlier cycles left to the analyst. The nonresponse problem is not academic. NCHS reports that in August 2021 to August 2023 the interview response rate was 34.6 percent and the examination response rate 25.7 percent, compared with 51.0 and 46.9 percent in the 2017 to March 2020 pre-pandemic period, and far higher rates in the early 2000s. NCHS's nonresponse bias assessment found no evidence that the sampled counties failed to represent the national population and concluded that the weighting adequately addressed the nonresponse it could detect. That is reassuring, and it is also exactly what a nonresponse analysis can conclude: it can check characteristics that are known for nonrespondents, not the health outcomes that were never measured in them. Analysts using the most recent cycle should read the assessment and consider, for their own outcome, whether nonresponse is likely to be related to it beyond what the weighting variables capture. Choosing the right weight The rule for choosing among NHANES weights is simple to state: use the weight associated with the smallest subsample that contributes a variable to the analysis. If an analysis uses interview variables and examination variables, the examination weight is correct, because only examined participants have all the variables. If it adds a laboratory variable measured on a fasting subsample, the fasting subsample weight is correct. If it adds dietary intake from two days of recall, the two-day dietary weight is correct, unless it also uses a fasting laboratory variable, in which case there is no official weight for the intersection, and NCHS's tutorials advise using the weight for the smaller subsample and accepting a modest approximation. Table 3 sets out the choice for common combinations. Table 3. Choosing an NHANES weight for common analyses (source: NCHS NHANES weighting tutorial). Variables in the analysis Weight to use Why Household interview only Interview weight All interviewed participants have the data Interview plus examination Examination weight Only examined participants have both Plus a laboratory test on a subsample Subsample weight for that test Only the subsample has the test Plus fasting laboratory values Fasting subsample weight Fasting subsample is smaller still Plus two days of dietary recall Two-day dietary weight Only two-day completers have the data Choosing the wrong weight is one of the most common errors in published NHANES analyses. Using the interview weight for an analysis of examination data gives weights that do not account for examination nonresponse and do not sum to the population total over the examined sample. Using the examination weight for a fasting glucose analysis ignores the fact that the fasting subsample was selected at random from examinees, with its own weight to reflect that selection. The error rarely changes a point estimate dramatically, but it is avoidable, and reviewers who know the survey will catch it. Combining cycles requires a further adjustment. The weights in each two-year cycle sum to the population total for that period. When two cycles are pooled, the combined weights would sum to twice the population, so each two-year weight is divided by the number of cycles pooled. NCHS provides special four-year weights for 1999 to 2002, because those two cycles were weighted to different census bases, and pre-pandemic weights for the combined 2017 to March 2020 file, because the partial 2019 to 2020 cycle cannot stand alone. Pooling 2021 to 2023 with earlier cycles is possible arithmetically but NCHS advises caution, because the interruption in data collection from April 2020 to July 2021 coincided with large changes in health care use, employment and schooling, and because the design changed. A pooled estimate spanning the pandemic gap does not describe any real period. Clusters, strata and the variance problem Weights fix the point estimate. They do not fix the standard error, and on their own they can make it worse. The standard error of a survey estimate depends on the design, and two features of the design matter most. Clustering inflates variance. People in the same county, the same neighbourhood and the same household resemble one another in health, diet, income and exposure. A sample of 5,000 people drawn from 15 counties contains less independent information than a simple random sample of 5,000 from the whole country, because much of the variation between the sampled people is shared within clusters. Leslie Kish's classic approximation expresses this through the design effect: for clusters of average size m and an intraclass correlation ρ, the variance is multiplied by roughly 1 + (m − 1)ρ. Even a small intraclass correlation produces a large design effect when clusters are large. With 300 examinees per county and ρ of 0.02, the design effect is about 7, and the effective sample size is a seventh of the nominal one. Unequal weighting also inflates variance, independently of clustering. Kish's approximation for this component is 1 + CV², where CV is the coefficient of variation of the weights. Weights that vary widely, as they do when some groups are heavily oversampled, reduce the precision of estimates for the whole population even as they increase precision for the oversampled groups. Stratification, by contrast, usually reduces variance slightly, because it guarantees that each stratum is represented in its expected proportion and removes between-stratum variation from the sampling error. To let analysts compute correct standard errors without revealing the actual counties sampled, NCHS releases masked variance units: a pseudo-stratum variable, SDMVSTRA, and a pseudo-PSU variable, SDMVPSU, constructed so that the variance structure is preserved but the geography is not identifiable. The design is represented with two pseudo-PSUs per pseudo-stratum. The number of degrees of freedom available for variance estimation is approximately the number of PSUs minus the number of strata, which in a single cycle is around 15, not the thousands implied by the number of participants. That low figure has consequences: confidence intervals should use t rather than normal critical values, and models with many parameters can exhaust the degrees of freedom available for testing them jointly. Estimating variance correctly Three families of method produce design-based standard errors, and the major statistical packages implement all of them. Taylor series linearisation approximates a nonlinear statistic, such as a ratio or a regression coefficient, by a linear function of weighted totals, and computes the variance of that linear function from the between-PSU variation within strata. It is the default in most software and the method NCHS generally recommends for NHANES. In R, the survey package, written by Thomas Lumley, specifies the design with a call along the lines of svydesign(ids = ~SDMVPSU, strata = ~SDMVSTRA, weights = ~WTMEC2YR, nest = TRUE, data = nhanes), after which functions such as svymean and svyglm produce linearised standard errors. In Stata, the equivalent declaration is svyset with the PSU, the weight and the stratum, followed by commands prefixed with svy. In SAS, the SURVEYMEANS, SURVEYFREQ, SURVEYREG and SURVEYLOGISTIC procedures take STRATA, CLUSTER and WEIGHT statements. Replication methods, including balanced repeated replication, the jackknife and the bootstrap, estimate variance by recomputing the statistic many times with modified weights and measuring how much it varies. Many surveys distribute replicate weights instead of design variables, because replicate weights can encode the design without revealing the structure. The American Community Survey public-use files, available through IPUMS, include 80 replicate weights computed by the successive difference replication method; the variance of an estimate is four eightieths of the sum of the squared differences between each replicate estimate and the full-sample estimate. The Current Population Survey's Annual Social and Economic Supplement supplies 160 replicate weights. An analyst using these files who ignores the replicate weights and computes naive standard errors will usually understate uncertainty, sometimes severely for small geographic areas. Generalised variance functions are simpler approximations, sometimes published by survey agencies, that relate the standard error of an estimate to its size. They are useful for quick checks but should not be used for formal analysis when design variables or replicate weights are available. Subpopulations: the most common error The single most frequent technical error in published survey analyses is to handle a subpopulation by deleting everyone outside it before declaring the design. An analyst studying women aged 20 to 44 drops all other participants from the file, specifies the survey design on what remains, and computes estimates. The point estimates are correct. The standard errors are not, because the number of sampled people in the subpopulation within each PSU is itself a random quantity that the variance calculation must account for, and because deleting records can remove entire PSUs from some strata, which breaks the variance computation or silently changes it. The correct approach is to declare the design on the full file and then restrict the analysis. In R's survey package this is done with the subset function applied to the design object, which retains the full design information. In Stata it is the subpop option to svy. In SAS it is the DOMAIN statement. The difference in standard errors between the correct and incorrect approaches is often small, but it is occasionally large, particularly for small subpopulations concentrated in a few PSUs, and there is no reason to accept the risk. A related issue is reliability. For small subgroups, an estimate may have so few observations or degrees of freedom that it should not be reported. NCHS publishes data presentation standards for proportions, which combine thresholds on effective sample size, degrees of freedom and the width of the confidence interval to decide whether an estimate is reliable enough to present. Adopting them, or a similar explicit rule, prevents the temptation to report every subgroup the software can produce. Reproducing the official figures The chapter began by describing a survey as a distorted picture with instructions for undoing the distortion. The best evidence that you have followed the instructions correctly is that you can reproduce an estimate NCHS has published. The NCHS Data Brief released with the August 2021 to August 2023 data reported the prevalence of obesity among adults, with a standard error. Reproducing that figure requires the correct file, the correct age restriction declared as a subpopulation, the correct examination weight, the correct design variables and, if the published figure is age-adjusted, the correct standard population. If your figure matches, your pipeline is sound. If it does not, the discrepancy will usually point directly at the mistake. Few checks in applied statistics offer so much protection for so little effort. Hashtags: #SecondaryDataAnalysis #PublicData #AdministrativeData #PopulationRegistries #CensusMicrodata #LongitudinalBiobanks #NHANES #IPUMS #UKBiobank #NordicRegistries #ComplexSurveyDesign #SurveyWeights #StratifiedSampling #ClusterSampling #VarianceEstimation #VolunteerSelectionBias #HealthyVolunteerBias #RecordLinkage #ProbabilisticLinkage #DataGovernance #TrustedResearchEnvironments #UnmeasuredConfounding #SensitivityAnalysis #SecondaryDataReproducibility #FutureOfSecondaryDataAnalysis

  • Scientific Visualization (Transforming High-Dimensional Data into Meaningful Graphics)

    Download the Book (PDF): Introduction A scientific figure is a measuring instrument pointed at a reader. That claim sounds odd at first, because we are used to thinking of instruments as things that point at the world. A spectrometer measures light; a sequencer measures nucleotides; a thermocouple measures temperature. But the last step of any measurement chain is a person forming a belief, and the graphic is the part of the apparatus that does that conversion. It takes numbers, which the human nervous system cannot apprehend directly in any quantity, and converts them into stimuli that a visual system with well-documented properties will process in well-documented ways. The output of the instrument is not ink. The output is an estimate in someone's head — of a difference, a trend, a cluster, an outlier, an effect size — together with a level of confidence in that estimate. Instruments can be calibrated or uncalibrated, precise or noisy, linear or distorting. So can graphics. A bar chart whose baseline has been cut at 90 rather than zero is not a neutral display with a caveat attached; it is an instrument with a systematic gain error, and it will produce biased readings in most people who use it, including people who know about the error. A rainbow colormap laid over a continuous field is not a stylistic preference; it is an instrument that inserts false boundaries at the cyan and yellow transitions and flattens real structure in the green, and it will make readers see edges that the data does not contain. A two-dimensional embedding of a fifty-dimensional dataset is not a picture of the data; it is a lossy transform with characteristic artifacts, and a reader who does not know which properties it preserves will confidently read off quantities it never encoded. This booklet is about designing that instrument deliberately. It rests on a single controlling idea, which is worth stating plainly at the outset because everything else follows from it: Every graphical choice is a bet about what a particular perceptual and cognitive system will do with a particular stimulus, and graphical excellence consists in making those bets knowingly, in favour of the signal that is actually in the data rather than the one the designer hopes is there. Two consequences run through the chapters that follow. The first is that a large part of figure design is not a matter of taste. There is a substantial experimental literature — running from Cleveland and McGill's psychophysical studies in the 1980s through crowdsourced replications and modern work on colour difference, correlation perception and uncertainty displays — that tells us, with reasonable confidence, how accurately people extract quantities from different visual encodings. Position along a common scale beats length, which beats angle, which beats area, which beats colour saturation. That ordering is not a style guide. It is an empirical result about error rates, and it has the same status in figure design that a detector's quantum efficiency has in instrument design. Disregarding it is allowed, but it costs accuracy, and the cost should be paid on purpose. The second consequence is less comfortable. If a figure is an instrument, then distorting it is not a design flaw but a measurement error, and the distinction between an honest mistake and a manipulation collapses at the point of use. The reader of a truncated axis is misled whether or not the author intended it. This is why the middle chapters of this book spend so much time on arithmetic — on what happens to perceived ratios when you map a quantity to the radius of a circle rather than its area, on what an aspect ratio does to a perceived slope, on what a bar chart of group means conceals about the distributions underneath it. These are not ethical questions in the first instance. They are questions about transfer functions. The ethics arrive only when you know the transfer function and publish anyway. Who this is for, and what it assumes The intended reader is someone who makes figures for a technical audience: a researcher preparing a paper, an analyst building a dashboard that other people will make decisions from, a data scientist reporting on a model, an editor deciding whether a submitted graphic supports the claim in the abstract. It assumes you can already produce plots in some tool — R with ggplot2, Python with matplotlib or one of its descendants, Julia, MATLAB, D3, Vega-Lite, or a commercial dashboard platform — and that the obstacle is not syntax. It does not assume any background in vision science, and it does not require one. The perceptual results that matter for this work are few in number, robust, and statable in ordinary language. Nor does it assume a particular discipline. The examples run across genomics, climate science, epidemiology, machine learning, physics and the social sciences, because the perceptual constraints are identical everywhere and only the conventions differ. One assumption is worth flagging because it shapes the whole treatment: this book is about figures that are supposed to be read quantitatively, by people who will act on what they read. That excludes a good deal of what gets called data visualisation. Graphics designed primarily to attract attention, to be memorable, or to be shared operate under different constraints and are optimised against different criteria; the evidence on their design is real but it is a different evidence base. A figure in a methods paper and an infographic in a magazine are both visualisations in the way that a caliper and a sculpture are both metal. What is deliberately left out Short books are made good by what they refuse. This one refuses three things. It does not teach any particular software. Tools change every few years; the constraints of the human visual system have not changed in the period covered by the written record, and will not change during the working life of anyone reading this. A book organised around a plotting library ages into a manual for a version nobody runs. The occasional code fragment appears where a concrete idiom is clearer than a paragraph, but the argument is tool-agnostic and should transfer to whatever you use. It does not attempt a survey of visualisation techniques. There are excellent taxonomies that catalogue treemaps, chord diagrams, hive plots, Sankey layouts and the rest; consulting one when you have an unusual data structure is sensible. But a catalogue teaches recognition, not judgement, and the failure mode in scientific publishing is almost never that the author did not know about parallel sets. It is that a perfectly ordinary scatterplot was drawn in a way that hid the thing it was supposed to show. It does not contain a single picture. That is a constraint worth naming in a book about graphics, and it is an instructive one. Everything here is carried by prose, with a handful of tables where the material is genuinely tabular. If the argument survives that treatment it is because the underlying claims are propositional — about what the eye does, about what an encoding costs, about what a transform preserves — and propositional claims can be written down. The discipline of explaining a design in sentences also happens to be an excellent test of whether you understand it. An author who cannot describe, step by step and without gesturing, why a particular panel arrangement makes a comparison easy, usually has not worked out why it does. The shape of the argument The book moves from the eye outward. The first chapter establishes what a scientific graphic is for, historically and functionally, and why the answer has consequences. The second describes the perceptual machinery the graphic is aimed at: what is extracted instantly and in parallel, what requires effortful serial attention, what is compressed logarithmically, and what is simply not encoded at all. The third turns those constraints into an ordering of visual encodings by measured accuracy, and shows how to choose among them when accuracy is not the only criterion. Chapters four and five treat the two areas where scientific figures most often go quietly wrong: colour, which is misused in more published figures than any other channel, and scales, where a handful of arithmetic decisions — baseline, transform, aspect ratio, binning, area mapping — silently determine what the reader concludes. Chapter six takes on distributions, and argues that the most common summary display in the life sciences is also among the least defensible. Chapters seven through nine take on the harder cases named in the title. Chapter seven is about high dimensionality: what scatterplot matrices, parallel coordinates, small multiples and tours can do, what dimension reduction methods such as PCA, t-SNE and UMAP preserve and destroy, and how to read an embedding without over-reading it. Chapter eight is about the craft of the publication-grade figure — panel layout, typography, direct labelling, captions, file formats, reproducible pipelines, accessibility — the unglamorous work that separates a figure that survives peer review from one that survives use. Chapter nine is about interaction and dashboards, where the designer's power over the reader is greatest, the feedback is weakest, and the discipline required is correspondingly higher. The conclusion is not a summary. It argues something the preceding chapters make possible but do not state: that the responsibility for a figure's accuracy cannot be discharged by disclosure, that the reviewer's role in this is badly underdeveloped, and that the most useful habit an individual researcher can adopt is to treat every figure as a claim about what the reader will conclude, and to check that claim on an actual reader before publishing. There is no shortage of advice about visualisation. What there is a shortage of is the connective tissue between the psychophysics and the plotting call — the reasoning that gets you from a result about angular judgement to a decision about whether this particular comparison should be a grouped bar chart or a slope graph. That connective tissue is what this book tries to supply. It will not make every figure obvious. It should make every figure deliberate. Chapter 1: The Figure as Instrument In September 1854, John Snow marked the deaths from cholera in the Soho district of London on a street map, one short black bar for each fatal case, stacked outward from the address where the person had lived. The result is often described as the birth of epidemiological mapping, and it is usually reproduced as an image of triumphant clarity: a dense black thicket around the Broad Street pump, thinning with distance, with the Marlborough Street workhouse and the Lion Brewery sitting as conspicuously empty islands in the middle of the cluster. What makes the map an argument rather than a decoration is not the marks. It is the correspondence between a spatial question and a spatial encoding. Snow's hypothesis was about proximity to a water source. Proximity is a property of position, and the map encodes position as position. No translation is required of the reader, no mental arithmetic, no holding of one number in memory while extracting a second. The pattern that would confirm or refute the hypothesis is the pattern the eye is best at seeing: a spatially contiguous concentration of marks. The brewery mattered because its workers drank beer from the company's own well and ale allowance rather than the Broad Street pump, so its emptiness was exactly the anomaly the hypothesis predicted. A table of the same deaths by address would have contained identical information and been effectively unreadable. That is the whole business, compressed into one historical example. A scientific figure works when the structure the analyst cares about is mapped onto a visual property the reader's perceptual system extracts accurately and without effort, and when nothing else in the display competes for that extraction. It fails in one of three ways: the structure is mapped onto a property the eye handles badly, the mapping is distorting, or the display is so cluttered that the relevant property cannot be isolated. Almost every specific piece of advice in the chapters that follow is an instance of one of those three failures and how to avoid it. Three jobs, three different figures It helps to be explicit about what a given graphic is for, because the same data supports very different displays depending on the job, and much bad figure design comes from building for one job and using for another. Exploration. Here the analyst is the reader, the audience is one person, the figure is disposable, and the goal is to find out what is in the data — including things nobody suspected. Speed and coverage matter more than polish. The right instinct is to make many plots quickly, to look at raw values rather than summaries, and to accept overplotting, ugly defaults and unlabelled axes as the price of iteration. John Tukey, whose Exploratory Data Analysis is still the best single book on this mode, described the purpose as forcing yourself to notice what you did not expect. The failure mode of exploratory graphics is premature tidiness: the analyst makes one beautiful plot of the thing they already believed and stops. Confirmation and diagnosis. A narrower job: checking whether an assumption holds. Residuals against fitted values, a quantile-quantile plot against a theoretical distribution, a trace plot from a sampler, a calibration curve. These figures exist to make a specific class of deviation visible, and they are effective in proportion to how boring they look when nothing is wrong. A good diagnostic plot has a null appearance — a structureless band, a straight line, a fuzzy caterpillar — so that departure is detected as a violation of an expected pattern rather than as a judgement about degree. Communication. The published figure, the one in the paper or the report or the dashboard. Now the reader is someone else, usually in a hurry, often reading only the abstract and the figures, and frequently less invested in the result than the author. The figure has to be self-contained, correctly labelled, robust to being viewed at reduced size on a screen, and honest under the least charitable reading. This is the mode this book is mainly concerned with, and the mode where the perceptual results bite hardest, because the author cannot stand next to the reader and explain. The distinction matters because the habits of one mode are actively harmful in another. Exploratory instincts — throw everything on, use whatever colours, trust the viewer to interrogate — produce dense, unreadable published figures. Publication instincts — summarise, smooth, aggregate, polish — produce exploratory analyses that never discover anything, because everything surprising was averaged away before it could be seen. Andrew Gelman and Antony Unwin made a related point about the gulf between statistical graphics and information visualisation: the two communities optimise for different things, and a display that wins design awards may be inert as an analytical instrument, while a display that reveals structure to a statistician may be impenetrable to everyone else. Neither is wrong. They are answers to different questions. What "excellence" meant, and what it means now The modern conversation about graphical quality begins with two figures from the nineteenth century and one book from the twentieth. Charles Joseph Minard's 1869 chart of Napoleon's Russian campaign encodes six variables — the army's position in two dimensions, its direction of travel, its dwindling size, the date, and the temperature during the retreat — in a single flowing band whose width shrinks from a broad ribbon at the Niemen to a thread at the return. Florence Nightingale's polar area diagrams of mortality in the Crimea, published in 1858, were built to make one comparison unmissable to a government audience: that deaths from preventable disease vastly exceeded deaths from wounds. Both are cited endlessly, and both are worth examining not as objects of admiration but as engineering. Minard's band works because quantity is mapped to width, a length-like judgement, along a path the reader traverses in one direction. Nightingale's works less well on strictly perceptual grounds — area judgements are poor, and her wedges vary in radius so that area grows as the square of the encoded quantity — but it succeeded rhetorically, which was the job it was built for. Edward Tufte's The Visual Display of Quantitative Information, published in 1983, turned scattered craft knowledge into a set of principles: maximise the proportion of ink devoted to data, remove redundancy, avoid decoration that carries no information, and never let the graphic's size effects exceed the data's. His "lie factor" — the ratio of the size of the effect shown in the graphic to the size of the effect in the data — remains the single most useful diagnostic in the field, precisely because it is a number you can compute rather than a matter of taste. Tufte's prescriptions have been challenged in detail. Experiments on so-called chartjunk have found that embellished charts are sometimes remembered better than minimal ones, and that memorability and accuracy are not the same target; work on the memorability of visualisations has found that unusual, pictorial and dense displays stick in memory even when they are not the most efficient to read. Maximal data-ink ratios can produce sparse, hard-to-parse graphics. But the challenges are refinements of a framework that was broadly right, and the core of it — that a graphic's job is to represent quantity faithfully and that ornament competes with that job — has survived four decades of testing better than most normative advice in any field. What has changed is the evidentiary basis. Tufte argued from example and from a strong aesthetic. The work that followed, beginning with William Cleveland and Robert McGill's experiments in the mid-1980s, argued from measured error rates in controlled tasks. That shift is the reason this book is possible. We are no longer restricted to saying that a pie chart is inelegant; we can say that angular judgements produce larger errors than length judgements by a measurable margin, and that the penalty grows as the number of slices grows. A 2021 review by Steven Franconeri and colleagues in Psychological Science in the Public Interest collected what this literature now supports with reasonable confidence, and it is a good indication of how much of figure design has moved from opinion to evidence. The reader is not a general-purpose decoder The most common unstated assumption in bad figures is that the reader will do arithmetic. They will not. A stacked bar chart with four categories asks the reader to compare the second segment across bars. Only the bottom segment sits on a common baseline; every segment above it starts at a different offset, so comparing them requires either mentally subtracting two endpoints or comparing lengths floating in space. Readers do neither reliably. A dual-axis chart asks the reader to hold two different scales in mind and to treat crossings of the two lines as meaningless, which is precisely what nobody does — the crossing looks like an event, because visually it is one. A log-scaled axis asks the reader to interpret equal vertical distances as equal ratios, which trained readers can do with effort and untrained readers routinely do not do at all. Every one of these designs is defensible in some circumstance. The point is that each imposes a cognitive tax, and that the tax is paid in accuracy and in errors that the reader does not know they are making. There is a well-documented illusion, sometimes called the bar-tip limit error, in which readers treat the region under the top of a bar as more likely than the region above it even when the bar's top represents a mean with symmetric uncertainty — the bounded shape of the bar leaks into the inference. Readers do not interpret marks; they perceive them, and then rationalise. The practical consequence is a design rule that will recur throughout this book: put the comparison you want the reader to make into a single, direct perceptual judgement, and make everything else subordinate. If the claim is that treatment A exceeds treatment B, the figure should be built so that the A–B difference is the most salient length or position difference on the page. If the claim is that a trend reversed in 2014, the reversal should be visible as a change in slope, not inferred by comparing two numbers from a table-like grid. A figure that supports six comparisons equally well supports none of them well. Figures as evidence, and what that implies In most empirical fields the figure is not an illustration of the evidence. It is the evidence, in the sense that it is the form in which almost all readers encounter the result. Reviewers examine figures before they examine methods. Readers who skim read abstract, figures and captions. Meta-analysts extract effect sizes from plots when tables are unavailable. In some fields the figure is the only place the raw data appears at all. That status carries obligations that are easy to state and often ignored. A published figure should show the data, not only a model of it. When sample sizes are small — and in much of experimental biology they are very small — the individual observations should appear, because the summary cannot be trusted to represent them and because the reader has a right to see how many points there are and how they are spread. When sample sizes are large, showing everything becomes counterproductive and a principled summary is required, but then the summary should be one whose failure modes are known. A published figure should make its uncertainty visible, and should say in the caption exactly what the interval represents. An error bar with no definition is uninterpretable: standard deviation, standard error and confidence interval differ by factors that depend on sample size, and readers who are not told which one they are looking at will assume whichever supports the conclusion they are inclined toward. A published figure should be readable by the population that will read it. Roughly eight per cent of men of northern European ancestry have some form of red–green colour vision deficiency, which means that in any large readership a substantial number of people cannot distinguish the two most commonly paired categorical colours. A figure that encodes its key contrast in red versus green is not merely suboptimal; for those readers it does not work at all. And a published figure should be reproducible from data by a script. This is partly a reliability argument — hand-edited figures acquire errors that nobody can trace — and partly an argument about revision. A figure that takes twenty minutes to regenerate will be regenerated when a reviewer asks; one that took a day of manual work in a drawing program will be defended instead. A failure with a known cost On the evening of 27 January 1986, engineers at Morton Thiokol argued with NASA managers about whether to launch the Space Shuttle Challenger the following morning, in forecast temperatures far below any previous launch. The engineers believed that cold degraded the resilience of the rubber O-rings sealing the joints of the solid rocket boosters. They had data: post-flight inspections from twenty-four previous shuttle flights, recording erosion and blow-by of the seals, along with the ambient temperature at each launch. The material faxed to Kennedy Space Center that night did not put those two variables against each other. It presented damage histories organised by flight and by joint, with temperatures mentioned in text, and it concentrated on the flights that had shown damage. Flights that had flown without incident — most of them warm — were largely absent from the presentation. The consequence is easy to state in the vocabulary of this chapter: the analysis was about a relationship between two quantitative variables, and the display encoded neither of them as position on a common scale. There was no plot of damage against temperature. Without the undamaged flights, there was not even a complete sample to plot. A scatterplot of an O-ring distress index against launch temperature, including every flight, shows a pattern that later statistical reanalysis confirmed: the incidence and severity of damage rise as temperature falls, and the forecast launch temperature lay far outside the range of prior experience. Whether such a plot would have changed the decision that night is not knowable, and it would be glib to claim it. Organisational pressure, a reversed burden of proof and a compressed timeline all contributed, and historians of the episode have rightly resisted the tidy moral that better graphics would have saved the crew. But the narrower claim survives the caveats, and it is the claim this book cares about. The engineers held a conviction that the data supported, and the form in which they presented it made that support invisible — invisible to the managers, and arguably to the engineers themselves, who could not point to the pattern because they had never drawn it. The general lesson is not that graphics are decisive. It is that the display determines which patterns are available for anyone in the room to reason about. A relationship that is never encoded as a relationship cannot be seen, argued over, or defended, and the people who own the data are not immune to that limitation merely because they know the numbers. That episode also illustrates something about selection. Restricting a display to cases where the outcome occurred is a sampling decision dressed as a formatting decision, and it is astonishingly common: the plot of responders only, the trajectory panel showing the cells that survived, the map of locations where the species was found with no marks where it was looked for and not found. The graphic inherits every bias in the subset it draws, and adds the authority of having been drawn. The chain from number to belief It is worth laying out the full chain the instrument metaphor implies, because each link is a place where the signal can be corrupted, and the rest of the book works through them in order. A quantity in the data is mapped to a visual variable — position, length, angle, area, colour, texture, motion. That mapping has a transfer function, which may be linear, compressive, or ambiguous. The visual variable is rendered into a stimulus, subject to the constraints of the medium: resolution, contrast, size, the printed page or the calibrated monitor or the projector in a bright room. The stimulus falls on a retina with extremely non-uniform spatial sampling, and is processed by early visual mechanisms that extract some properties in parallel across the whole field and others only within the narrow window of attention. The extracted properties are combined with prior expectations, with the caption, with the conventions of the discipline, and with whatever the reader wanted to be true. Out comes a belief. Corruption can enter at any link. A distorting mapping — quantity to radius rather than area — corrupts at the first. A rainbow colormap corrupts at the third and fourth, by making the perceptual distance between adjacent values wildly non-uniform. Overplotting corrupts at the fourth, by making density illegible where it matters most. An unlabelled interval corrupts at the fifth. The instrument fails as a whole if any one link fails, and the author is responsible for all of them, because the reader has no way to inspect the chain. This is why the next chapter starts with the retina rather than with plot types. You cannot design an instrument without knowing the characteristics of the detector at the far end, and in scientific visualisation the detector is a visual system whose properties are known well enough to design against. Chapter 2: What the Eye Actually Does The visual system is often described as though it were a camera feeding a screen inside the head. Nothing about it works that way, and almost every useful design rule in visualisation follows from the ways it does not. Four properties matter most for figure design. The retina samples space with extreme non-uniformity. Certain visual features are extracted across the entire field at once, while others require attention to be aimed at one location at a time. Magnitude estimation is systematically non-linear, and the non-linearity differs by feature. And visual memory across glances is far poorer than introspection suggests. Each of these has direct consequences for how a graphic should be built, and the consequences are often the opposite of what a designer's intuition proposes. Acuity is a narrow spotlight Detailed vision is confined to the fovea, a region of retina subtending roughly two degrees of visual angle — about the width of a thumbnail at arm's length, or a coin held at a metre. Cone density falls off sharply outside it, and spatial resolution falls with it. At ten degrees of eccentricity, acuity is a small fraction of what it is at the centre. Crowding compounds the loss: in peripheral vision, nearby objects interfere with one another's identification even when each is individually resolvable, so a dense scatter of small marks off to the side is not merely blurry but structurally unreadable. What this means in practice is that reading a figure is a sequence of fixations. The eye lands, extracts detail from a small region, and jumps — three or four times a second, with vision largely suppressed during the jump. A complex multi-panel figure is not apprehended; it is scanned, in an order determined partly by layout and partly by what the periphery flags as worth looking at. Two design consequences follow. First, anything that must be read precisely — a tick label, a small annotation, a distinction between two similar symbols — has to be at the point of fixation, which means the figure's design should control where fixations go. Direct labelling of a line at its end works because the label is where the eye already is after tracing the line; a legend in a corner works less well because each lookup costs a saccade, a fixation, a return, and a memory operation in between. Second, peripheral vision is not useless — it is what guides the next fixation — but it only carries coarse information: large-scale luminance structure, strong colour blobs, motion, orientation at low spatial frequency. A figure whose important structure is visible only at high spatial frequency will never attract the eye to it. This is the honest argument for a degree of visual salience in scientific figures. Not decoration, but the deliberate use of size, contrast and position to make the eye land where the evidence is. If the key comparison is between two of twelve lines, those two should be heavier and darker and the others grey, because otherwise the reader's fixation sequence is determined by accident. Preattentive features: what is free Some visual properties are computed in parallel across the whole visual field, early and fast, before attention is deployed. The classic demonstration is visual search: a red dot among blue dots is found in roughly the same time whether there are ten distractors or a hundred, because the colour difference is available everywhere at once. Searching for a red square among red circles and blue squares — a conjunction of two features — takes time proportional to the number of items, because conjunctions require attention to bind features at each location in turn. Anne Treisman and Garry Gelade's feature integration theory, published in 1980, formalised this distinction, and while the theory has been revised repeatedly since, the empirical pattern is solid and directly useful. The features that support parallel search include hue, luminance, size, orientation, curvature, motion, flicker, spatial position, length, width, and a handful of others including stereoscopic depth and certain texture statistics. A rough working list for figure design would be: colour, brightness, size, orientation, shape at a coarse level, position, and motion. Three practical rules come out of this. One channel per question. If you want an outlier class to pop out, encode it on a single preattentive dimension that nothing else uses. If group membership is encoded by colour and the same colour is also doing duty for measurement condition, neither pops. Conjunctions are expensive. A figure that requires the reader to find "the filled triangles that are also blue" imposes a serial search. Sometimes that is unavoidable — a scatterplot with two categorical variables has to encode both — but it should be a conscious cost, and the more important variable should get the stronger channel. Pop-out is graded, not binary. The speed advantage depends on how different the target is from the distractors relative to the variability among the distractors. A category encoded in a slightly darker blue among other blues does not pop out at all. Designers routinely underestimate how large a difference has to be to work. There is a corresponding negative rule. Because pop-out is automatic, anything that is highly salient will be attended whether or not it is important. Gridlines at full contrast, heavy plot borders, saturated background panels, drop shadows, and bold reference lines all compete for the same early visual resources as the data. This is the defensible core of the data-ink argument: not that ink is inherently wasteful, but that salience is a fixed budget and non-data elements spend it. Magnitude perception is compressive and channel-specific Ask people to judge how much bigger one stimulus is than another and their answers are not proportional to the physical difference. For many continua the relationship follows a power law — Stevens's law — with an exponent characteristic of the dimension. Brightness and area are strongly compressive: doubling the physical quantity produces a perceived increase considerably less than double. Line length is nearly veridical, with an exponent close to one. This is why a bar chart works and a bubble chart does not, quite apart from the geometry. The geometry compounds the psychophysics in one particularly common error. If a quantity is mapped to the radius of a circle, the area grows as the square of the quantity, so a value three times larger occupies nine times the ink. Mapping to area instead removes that error, but leaves the perceptual compression: people underestimate area ratios, typically reading a circle of nine times the area as something like four or five times the value. The cartographic literature has grappled with this since the 1970s, when psychophysical scaling of proportional symbols led to proposals for deliberately exaggerated area scales to compensate. Those corrections are contested and rarely used in scientific plotting, which leaves the practical advice simple: use area encodings only for rough orders of magnitude, never for comparisons the argument depends on, and if you must use them, scale by area and label the extremes. A related result concerns the just-noticeable difference. Ronald Rensink and Gideon Baldridge showed that the precision with which people can discriminate correlation in a scatterplot follows a regular law, with discrimination becoming markedly finer as correlation approaches one; Lane Harrison and colleagues extended this across nine different chart types and found the same Weber-law form, with different sensitivity constants per chart type. The practical implication is under-appreciated: the difference between r = 0.2 and r = 0.3 is nearly invisible in a scatterplot, while the difference between r = 0.9 and r = 0.95 is easy to see. If your argument turns on a small difference in weak correlation, the plot will not carry it and you should report the number. Colour is three channels, and they are not equal Colour is not a single continuum. Human colour vision is trichromatic at the receptor level, but the signals are immediately recoded into one achromatic channel (light–dark) and two chromatic opponent channels (roughly red–green and blue–yellow). These channels have very different properties. The luminance channel carries high spatial resolution. Edges, texture, shape from shading and fine detail are all conveyed by luminance. The chromatic channels are low-pass: they carry much coarser spatial information, which is why image compression schemes throw away colour resolution first and why thin coloured lines on a background of similar lightness are hard to see even when the hue difference is large. This single fact explains a large fraction of the colour failures in published figures. A heatmap that encodes its quantity purely in hue, with roughly constant lightness, presents fine spatial structure on the channel least able to resolve it. A categorical palette in which all the colours have similar lightness produces lines that are perfectly distinguishable at large size and mushy at print size. A continuous colormap whose lightness is non-monotonic — rainbow being the archetype — creates apparent edges wherever lightness happens to reverse. Ordered data therefore needs ordered lightness. This is the one colour rule that resolves most disputes: if the variable has a magnitude, the colormap must change monotonically in perceived lightness across its range, with hue and saturation doing secondary work to widen the discriminable range. Chapter four develops this at length, but it is fundamentally a perceptual constraint rather than an aesthetic preference. Colour vision deficiency adds a second constraint. Anomalies of the red–green opponent channel affect around eight per cent of men and under one per cent of women of northern European descent, with somewhat different rates in other populations. For these readers, red and green of matched lightness are close to indistinguishable. The robust solutions are to vary lightness along with hue, to use blue–orange rather than red–green pairings, and to add a redundant non-colour channel — shape, line style, direct labels — for anything essential. Visual working memory is tiny Between one fixation and the next, very little is retained. The phenomenon of change blindness — a large change to a scene going unnoticed when it occurs during a saccade, a blink, or a brief blank — demonstrates that we do not hold a detailed internal image. Estimates of visual working memory capacity converge on something like three to four simple objects, and fewer when the objects are complex. Three implications for figures, all of them significant. Comparisons that require memory are unreliable. Any design that forces the reader to look at one panel, memorise a value, and look at another has inserted a low-capacity buffer into the measurement chain. This is why animation is frequently worse than small multiples for comparing states: animation puts the comparison in memory, while small multiples put it side by side where it becomes a simultaneous perceptual judgement. Experiments comparing animated and static displays of trend data have generally found small multiples more accurate, with animation sometimes preferred and enjoyed more even while producing more errors. Legends are memory operations. Every legend lookup is an encode–saccade–match cycle, and with more than about five or six categories the reader begins losing track. Direct labelling removes the cycle entirely, and is the single highest-yield change available in most line charts. Ordering is a memory aid. Sorting categories by value rather than alphabetically, arranging small multiples so that adjacent panels differ in one variable at a time, and keeping axis ranges identical across panels all reduce what the reader must hold in mind. Inconsistent axis ranges across panels of a multi-panel figure are the most common way to make a comparison impossible while appearing to enable it. Absolute judgement is severely limited There is a difference between telling two stimuli apart when they sit side by side and identifying one in isolation. The first is relative discrimination and it can be remarkably fine; the second is absolute judgement and it is poor. George Miller's 1956 review of the experimental literature found that across a wide range of unidimensional continua — tone pitch, loudness, line length, position of a marker on a line, saltiness of a solution — people could reliably assign stimuli to only about seven distinct categories, and often fewer. The number rises somewhat when several dimensions are combined, but never dramatically. This is why a legend that maps twelve shades of blue to twelve values is not a quantitative encoding. A reader can tell two adjacent patches apart when they touch; asked to look at a patch in the middle of a map and say which legend entry it matches, they cannot do it beyond a handful of levels. The same limit applies to line thickness, symbol size, texture density and hue. For colour, careful work on colour difference perception in visualisation — including models that account for the size of the mark, since small marks are much harder to discriminate than large ones — puts the number of reliably identifiable categorical colours at somewhere between eight and twelve under good conditions, and the number for small scattered points considerably lower. Two design rules follow. When a channel with low absolute-judgement capacity carries a quantitative variable, treat the encoding as ordinal — the reader will be able to say "more" or "less", not "how much" — and provide the actual numbers wherever the argument needs them. And when a categorical variable has more levels than the channel can carry, do not add colours; change the design. Facet into small multiples, group the minor categories into an "other" class, label the important series directly and grey out the rest, or accept that the figure is a texture showing overall shape rather than a lookup table. The corollary is that redundant encoding is not waste. Encoding the same variable twice — colour and position, colour and shape, colour and direct label — converts an absolute judgement into a relative one, and costs nothing but ink. In a scatterplot with four groups, using both colour and marker shape is often derided as redundant; in practice it is what makes the figure work when it is printed in greyscale, viewed by a reader with a colour vision deficiency, or reduced to a column width. Gestalt grouping: the eye organises before you do Early vision imposes organisation on marks whether or not it corresponds to structure in the data. Marks that are close together are grouped; marks that are similar in colour or shape are grouped; marks that fall along a smooth continuation are read as one object; a closed contour is read as a region with an inside; elements that move or change together are read as a unit. These principles are tools when exploited and traps when ignored. Grouping by proximity is why whitespace is the most effective panel separator and why a grouped bar chart with equal spacing between all bars fails to communicate its grouping. Continuity is why connecting points with a line asserts that the intervening values exist — a line between two categorical levels is a factual claim about interpolation, and usually a false one. Closure is why a filled area under a curve reads as a quantity, so that filling under a line whose baseline is not zero quietly implies a magnitude that is not there. Common fate is why brushing in an interactive display works so well: making the selected points change together binds them into an object across panels. The most common gestalt error in scientific figures is unintended grouping by colour. Using one palette for experimental conditions in panel A and re-using the same palette for a different variable in panel B makes the reader group across panels, because colour identity is a stronger grouping cue than panel membership. Within a single figure, a colour should mean one thing. Putting the constraints together Taken together, these properties describe a detector with an unusual specification: extremely high resolution in a tiny window, coarse but parallel processing everywhere else, accurate for position and length, compressive for area and brightness, split into one sharp achromatic channel and two blurry chromatic ones, with almost no memory between glances, and a strong built-in tendency to organise marks into groups and continuations. A figure designed for that detector looks like this. The critical quantity is encoded as position or length on a common, unbroken scale. The critical comparison is adjacent, so that no memory is required. Categories are few, and distinguished on more than one channel. Colour with a magnitude varies in lightness monotonically. Salience is spent on the data rather than on frames, grids and backgrounds. Labels sit next to the things they label. Panels share scales, and differ in one variable at a time. Nothing is asked of the reader that requires arithmetic, and nothing important is asked of peripheral vision. That specification is abstract. The next chapter makes it concrete by ranking the available encodings against measured human performance, which turns "encode the critical quantity well" into a decision procedure. Chapter 3: The Ranking of Encodings In 1984 William Cleveland and Robert McGill published a paper in the Journal of the American Statistical Association that did something the field had not done before: it treated the choice of graphical encoding as an empirical question with a measurable answer. Their method was to decompose graph reading into what they called elementary perceptual tasks. When a reader extracts a quantity from a graphic, they are performing one of a small number of operations — judging position along a common scale, judging position along identical but non-aligned scales, judging length, judging direction or angle, judging area, judging volume or curvature, judging shading or colour saturation. Different chart types demand different tasks. A dot plot requires position along a common scale. A stacked bar chart requires length. A pie chart requires angle. A bubble chart requires area. A choropleth map requires colour. Cleveland and McGill then ran experiments in which subjects estimated ratios between marked values in displays that isolated each task, and measured the error. The result was an ordering, from most accurate to least, which Table 1 sets out alongside the chart forms that depend on each task. Table 1. Elementary perceptual tasks ranked by measured accuracy, after Cleveland and McGill (1984), with the chart forms that rely on each. Rank Perceptual task Typical chart forms Use for 1 Position along a common scale Dot plot, scatterplot, line chart, aligned bars Quantities the argument depends on 2 Position on identical, non-aligned scales Small multiples with shared axes Quantities compared across panels 3 Length Stacked bars, floating bars, Gantt spans Secondary quantities 4 Angle and slope Pie charts, slope graphs, line angles Trend direction, not magnitude 5 Area Bubble charts, treemaps, cartograms Rough orders of magnitude 6 Volume, curvature Three-dimensional symbols Avoid 7 Shading, colour saturation Choropleths, heatmaps Ordinal patterns, spatial context The ordering has held up. Jeffrey Heer and Michael Bostock replicated the core experiments in 2010 using crowdsourced participants, obtaining results consistent with the original ranking, and extended the method to questions Cleveland and McGill had not addressed — including how the accuracy of rectangular-area judgements varies with the aspect ratio of the rectangles, and how small a chart can be made before accuracy degrades. Later work has filled in the constants for particular comparisons. The qualitative ordering is now about as settled as anything in the field. What the ranking is and is not It is a statement about accuracy in ratio estimation, and nothing else. It does not say that pie charts are forbidden, that heatmaps are bad, or that every figure should be a dot plot. It says that if a reader must extract a quantitative comparison, the encodings near the top of the list will produce smaller errors than the ones near the bottom, and it quantifies roughly by how much. Three qualifications keep it from being applied mechanically. Accuracy is not the only objective. A choropleth map encodes quantity in colour, the worst channel on the list. It is nevertheless the right display for many geographic questions, because the spatial arrangement is itself the point and no other encoding preserves it. A treemap uses area, ranked fifth, but it can show a thousand nested items in a space where a ranked dot plot would need thirty pages. The rule is not to choose the top encoding always; it is to know what you are paying when you go down the list, and to spend it on something. The task matters more than the chart type. A pie chart is poor for comparing two similar slices and perfectly adequate for the judgement "this is about half". A stacked bar is poor for comparing middle segments and fine for showing that the total is roughly constant. Naming a chart type good or bad without naming the task is a category error, and it is the reason chart-type arguments on the internet never resolve. Ratio estimation is not the only elementary task. Later work by Robert Amar, James Eagan and John Stasko catalogued the low-level operations analysts actually perform: retrieve a value, filter, compute a derived value, find extrema, sort, determine range, characterise distribution, find anomalies, cluster, correlate. Several of these are not ratio judgements at all. Finding anomalies benefits from preattentive pop-out; characterising a distribution benefits from a display that shows shape; correlating benefits from a scatterplot, which is not near the top of the ranking for reading individual values but is unrivalled for reading joint structure. Match the encoding to the operation, not to a list. From ranking to a design procedure The ranking becomes useful when combined with a second idea, due to Jock Mackinlay in a 1986 paper on automating graphic design. He proposed two criteria that any mapping from data to graphic must satisfy. Expressiveness: the graphic should express all the facts in the data, and only the facts in the data. A mapping that leaves out information fails; so does one that adds information the data does not contain. Plotting unordered categories along a positional axis violates the second half, because position implies an ordering that does not exist — which is why alphabetical category ordering is so pernicious, since it asserts a sequence that means nothing while suppressing the meaningful one. Connecting categorical values with a line asserts interpolation. Using a continuous colour ramp for nominal categories asserts that some categories are between others. Effectiveness: among expressive mappings, choose the one the reader decodes most accurately, which is where the Cleveland–McGill ranking enters, adjusted for the data type. The effective channels differ by type. Quantitative data is best on position, then length, then angle, then area, then density, then saturation, then hue. Ordinal data is best on position, then density, then saturation, then hue, then texture, then connection. Nominal data is best on position, then hue, then texture, then connection, then containment, then density, then shape. Together these give a procedure that will produce a defensible first draft for almost any dataset: 1. List the variables and their types: quantitative, ordinal, nominal, temporal, spatial. 1. Identify the one or two comparisons the figure exists to support. Write them as sentences. 2. Assign the variables in those comparisons to the highest-ranked channels available for their type. 3. Assign the remaining variables to the channels left over, or move them out of the figure into faceting, or drop them. 4. Check expressiveness: does the encoding imply any ordering, continuity, or magnitude that the data does not have? Step 2 is the one people skip, and it is the one that does the work. A figure whose purpose is "show the data" has no basis for choosing among encodings and generally ends up with whatever the plotting library does by default. A worked case: six treatments, four assays, three replicates Abstract rules are easy to agree with and hard to apply. Take a concrete and thoroughly ordinary dataset: a screening experiment with six compounds, each measured on four assays, with three biological replicates per combination. Seventy-two numbers. The question the paper asks is whether any compound shows activity across multiple assays rather than in one alone. The default output of most analysis pipelines is a heatmap: compounds on rows, assays on columns, mean activity in colour. It fits on a postage stamp and it is what everyone in the field draws. Run it through the procedure. The comparison the figure exists to support is: within a compound, is activity elevated on more than one assay, and is that elevation large relative to replicate noise? That comparison involves a quantitative variable (activity), two nominal variables (compound, assay), and a measure of dispersion. The heatmap puts the quantitative variable on colour — rank seven — and discards the replicates entirely. It is expressive about the two nominal variables and mute about the one thing the question turns on. A better mapping falls out almost mechanically. Activity is quantitative and carries the argument, so it goes on position along a common scale. Assay is nominal with four levels, and the comparison is within compound and across assay, so assay goes on the other positional axis — as a categorical axis, ordered not alphabetically but by something meaningful, such as mechanistic proximity or overall signal. Compound has six levels and the comparison across compounds is secondary, so it becomes the faceting variable: six small panels, shared axes, arranged in a single row or a two-by-three block. Replicates are plotted as individual points, because with n of three no summary is credible and three dots cost nothing. If a summary is wanted, a short horizontal line at the mean over the points serves, with no bar and no error bar pretending to be an interval. The result is six panels of twelve points each. Position along a common scale carries the quantity; panels on identical non-aligned scales carry the compound comparison, which is rank two; the reader sees replicate spread directly rather than inferring it. Nothing is encoded in colour at all, which means the figure survives greyscale printing and can spend colour later if a highlight is needed. Now the honest accounting. The heatmap has two real advantages that the faceted version gives up. It is far more compact, so if the screen had contained sixty compounds rather than six the faceted display would be unusable and the heatmap would be the only viable option. And a heatmap sorted by clustering reveals block structure — groups of compounds behaving alike — that separated panels do not, because similarity of colour pattern is a texture judgement made across the whole matrix at once. If the paper's claim were about clusters of compounds, the heatmap would be the right instrument and the faceted plot the wrong one. That is what applying the ranking actually looks like. Not "heatmaps are bad", but: name the comparison, find the channel that serves it, notice what the alternative buys, and choose with the cost in view. The same six compounds might appear as a heatmap in a supplementary screen-wide panel and as a faceted dot plot in the main figure that supports the central claim, and both would be correct. One further move is worth noting because it recurs. Suppose the question changes slightly: not "which compounds are active" but "did activity go up or down between a control and treated condition for each compound". Now the data is paired, and pairing is structure the display should preserve. A grouped bar chart of control and treated means destroys it — the reader sees two population summaries and cannot recover which control value belongs with which treated value. A slope graph, in which each compound is a line segment joining its control value on the left axis to its treated value on the right, encodes the paired change as slope and direction. Slope sits at rank four for magnitude, but the judgement here is not magnitude; it is sign and consistency, which slope and crossing patterns deliver preattentively. A dozen lines all tilting the same way is a result visible in a quarter of a second, and no bar chart of means conveys it at all. The general lesson is that the ranking answers the question "how accurately will the reader read a quantity from this channel", and many scientific claims are not about quantities read one at a time. They are about consistency, ordering, pairing, sign, dispersion, or the presence of an anomaly. For those, ask which display makes the relevant pattern a single perceptual object, and use the ranking to settle the residual choices. Position, and why it wins so decisively Position along a common scale wins for several converging reasons. It is preattentive, so gross patterns are available without search. It is nearly veridical — the psychophysical exponent for length and position judgements is close to one, so perceived differences track real differences. It supports both relative judgements and, with axis labels, approximate absolute ones. It permits many marks in a small space without the marks interfering. And it is the only channel that comfortably supports the judgement of two variables at once, which is what makes the scatterplot the most information-dense elementary display we have. The practical consequence is a strong default: when in doubt, put the quantity on an axis. An enormous fraction of figure improvement in practice consists of converting something else into a positional encoding. A pie chart of shares becomes a sorted dot plot. A stacked bar of multiple series becomes a set of small multiples of aligned bars, or a line chart if the categories are ordered. A bubble chart of three variables becomes a scatterplot of the two that matter with the third as a facet. A heatmap whose rows are being compared becomes a set of profile lines. A table of numbers whose point is a ranking becomes a dot plot, which Cleveland introduced for exactly this purpose and which remains badly underused. Bars deserve a note of their own. A bar chart encodes quantity in both position (the top of the bar) and length (the bar itself), which is why it reads so easily. The doubling is also why the baseline cannot be moved: truncating the axis breaks the length encoding while leaving the position encoding intact, and the reader perceives both. A line chart has no length encoding, which is precisely why its axis may legitimately be truncated when the data warrants it. This is the single most useful thing to know about the truncation debate, and it is developed further in chapter five. When to go down the list on purpose There are good reasons to use a lower-ranked encoding, and it is worth naming them so that the choice can be made rather than drifted into. Space and cardinality. With ten thousand items, position-based displays run out of room. Colour and area scale to densities that position does not. A genome-wide heatmap, a treemap of a filesystem, a hexbin density plot — these accept a weaker channel to gain capacity. Preserving a spatial or topological structure. Maps, brain images, microscopy, network layouts, and anything where the spatial arrangement of the data has physical meaning. Here position is already spoken for, and the measured variable must go somewhere else. Colour is the standard answer and the reason careful colormap design matters so much in these fields. Gist over precision. If the message is "most of this is one category" or "the pattern is roughly uniform", a lower-ranked encoding may deliver the gist faster. A pie chart genuinely does communicate part-of-whole membership quickly, which is why it survives. Convention. Fields have reading habits. Manhattan plots, Kaplan–Meier curves, volcano plots, forest plots, Bland–Altman plots, receiver operating characteristic curves: each is a compact convention that a trained reader parses far faster than a from-scratch alternative, even where the encoding is not theoretically optimal. Deviating from a strong convention imposes its own cost, and it should be done only when the conventional form actively misleads. What is not a good reason is novelty, or the desire to fit more variables into one display than the reader can extract. The commonest expression of the second is the three-dimensional bar chart, which converts an accurate position judgement into an inaccurate volume judgement, adds occlusion, and introduces perspective distortion that makes bars at the back systematically smaller. It has no defensible use in quantitative work. Redundancy, and the grammar underneath The encoding ranking presumes that each variable is mapped to one channel. In practice good figures often map one variable to two channels at once. A series encoded by both colour and line style survives greyscale printing. Points encoded by both colour and shape survive colour vision deficiency and small size. A sorted bar chart encodes value in both length and vertical position. Redundancy costs nothing except the channels it uses up, and it converts absolute judgements into easier ones. The only caution is that a redundantly encoded variable consumes channels that another variable might have needed, and that redundancy between different variables — where colour means condition in one place and timepoint in another — is the opposite of helpful. Underneath all of this sits a formal structure worth naming, because it changes how you think about plots. Leland Wilkinson's The Grammar of Graphics, published in 1999 and implemented in Hadley Wickham's ggplot2 and in Vega-Lite, treats a statistical graphic not as a chart type chosen from a menu but as a composition: data, a statistical transformation, a mapping from variables to aesthetic channels, geometric objects that realise those channels, a coordinate system, and scales with guides. Under that grammar, a pie chart is a stacked bar in polar coordinates, and a bar chart and a histogram differ only in the statistical transformation applied before drawing. Chart types stop being primitives. The practical value of the grammar is that it makes the encoding decision explicit and forces it into the code. When you write a mapping of a variable to the y position and another to colour, you have stated your design in the same vocabulary the perceptual literature uses, and the question "is this the right channel for this variable?" becomes a question you can actually answer. Hashtags: #ScientificVisualization #DataVisualization #HighDimensionalData #GraphicalPerception #VisualEncoding #PerceptualAccuracy #ClevelandMcGillRanking #PositionEncoding #ColourPerception #ScientificColormaps #DataIntegrity #UncertaintyVisualization #DistributionVisualization #SmallMultiples #DimensionalityReduction #PrincipalComponentAnalysis #TSNE #UMAP #ScatterplotMatrices #ParallelCoordinates #PublicationGraphics #ReproducibleFigures #VisualizationAccessibility #GrammarOfGraphics #FutureOfScientificVisualization

  • Scientific Storytelling and Public Engagement (Communicating Beyond Academic Walls)

    Download the Book (PDF): Introduction Most researchers learn to communicate in a single, highly specialised genre. The journal article has a fixed architecture: introduction, methods, results, discussion. It assumes a reader who already cares, who shares a vocabulary, who will forgive a slow opening because the payoff is a contribution to a literature they are tracking. It rewards caution, qualification and completeness. It withholds the conclusion until the evidence has been laid out, and it treats the author's personality as noise to be filtered out. These are virtues inside the walls of a discipline. Outside them, almost every one of those habits becomes a liability. A parent deciding whether to vaccinate a child, a city councillor weighing a flood defence scheme, a journalist with ninety minutes before deadline, a video editor asking for a thirty-second version of a four-year project, a stranger on social media who has decided your paper is part of a conspiracy: none of these people is the reader the journal article was designed for. They arrive with different questions, different stakes, different amounts of time and different reasons to trust or distrust you. Communicating with them is not a matter of saying the same thing more slowly or with fewer long words. It is a different task, with its own structure, its own failure modes and, increasingly, its own body of evidence. That last point matters. For a long time, advice on public communication came in two flavours. The first was inspirational: be passionate, tell stories, get out of the lab. The second was defensive: stay away from journalists, never speculate, let the work speak for itself. Neither was grounded in much more than anecdote. Over the past four decades, however, psychologists, risk researchers, political scientists and communication scholars have built a serious literature on how people actually take in, interpret and act upon scientific information. Baruch Fischhoff and his colleagues at Carnegie Mellon developed methods for finding out what people already believe before trying to change it. Michael Dahlstrom has reviewed what is known about how narrative works on non-expert audiences and where it becomes ethically fraught. Matthew Nisbet and others have shown how the frame around a scientific finding shapes which audiences listen and what they conclude. Roger Pielke Jr. has given researchers a vocabulary for the different roles they can play in policy debates. Randy Olson, a marine biologist turned filmmaker, distilled narrative structure into a template that working scientists can apply in an afternoon. Studies of press releases have traced where exaggeration in health news actually comes from, and the answer has made university press offices uncomfortable. Research on misinformation has overturned some confident claims about how corrections backfire. This booklet draws on that literature to give researchers a practical, evidence-informed approach to communicating with the public, with policymakers and with the media. Its controlling argument is simple to state and demanding to practise: communication beyond the academy succeeds when researchers start from what a particular audience needs to understand or decide, build a story that the evidence can honestly carry, and choose deliberately what role they are playing, so that the goal becomes earning trust rather than winning an argument. Each part of that sentence carries weight. Start from the audience. The dominant instinct among scientists is to begin with the finding and work outward. The evidence consistently points the other way. People interpret new information through mental models they already hold, through the values and identities of the groups they belong to, and through the decisions they are facing. A message designed without knowing any of this will be misunderstood in predictable ways. The first chapter examines why the old "deficit model", in which public scepticism is treated as a shortage of facts to be topped up, fails so persistently, and what better models ask communicators to do instead. Build a story the evidence can carry. Narrative is not decoration. It is how human beings organise cause, consequence and meaning, and it is remarkably effective at holding attention and aiding recall. It is also a powerful instrument that can outrun the data it describes. The second chapter lays out the structure of scientific narrative, from Olson's "And, But, Therefore" template to the larger arcs that suit longer formats, while taking seriously the ethical questions Dahlstrom and others have raised. The third chapter turns to framing: the choice of which aspect of a finding to put in the foreground, and why that choice so often determines who listens. The fourth deals with language and numbers, the level at which most communication actually breaks down, and offers concrete techniques for removing jargon and presenting risk in forms people can use. Choose your role. A researcher who speaks to a minister, a reporter or a public meeting is doing something different from one who publishes a paper, and the difference is not only one of style. It raises questions about authority, advocacy and responsibility. The fifth chapter uses Pielke's typology of roles, together with research on communicating uncertainty, to help researchers decide in advance how they will engage with policy. The sixth covers the practical realities of working with journalists, from writing an accurate press release to preparing for a live interview. The final two chapters address formats and hazards that the older advice barely anticipated. Video abstracts, short explainers and visual summaries are now routine requests from journals, funders and institutions, and they reward a distinct set of skills in scripting, pacing and design. Social media, meanwhile, has made researchers directly visible to very large audiences, which brings reach and conversation but also coordinated hostility, misinformation and, in the worst cases, threats. The seventh chapter treats the video abstract as a genre with its own rules. The eighth offers a way of distinguishing legitimate critique from bad-faith attack, and of responding to each without being either defensive or naive. A few words on what this booklet is not. It is not a guide to becoming a celebrity scientist, and it does not assume that every researcher should be doing public engagement all the time. Some people will never want to appear on camera; some fields rarely touch public controversy; some career stages leave little room for anything but the core work. The aim is to make the engagement you do choose to undertake more effective and less risky, and to help you recognise when a request is one you should decline. Nor is it a manual for persuasion in the marketing sense. Several of the techniques described here, narrative and framing above all, can be used to push audiences towards conclusions the evidence does not support. The booklet is explicit about where those lines fall, because a researcher's credibility is a long-term asset that a single overreaching message can spend. The examples throughout are drawn from real, documented cases wherever possible: the L'Aquila earthquake trial, the "arsenic life" controversy, the conflicts over the MMR vaccine, the harassment of researchers during the COVID-19 pandemic, and others. Where an example is invented to illustrate a technique, it is labelled as hypothetical. The research cited is real and listed in the Notes and Further Reading at the end, so that readers who want to go deeper can check the sources and draw their own conclusions. One final point about disposition. Researchers are trained to be sceptical of rhetoric and to trust data, and many find the language of "storytelling" and "messaging" faintly distasteful. That instinct is healthy and worth keeping. But the choice is never between communicating with a story and communicating without one. Every press release, every briefing, every conference talk already has a structure, a frame and a set of word choices; the only question is whether those were chosen deliberately or by default. A default structure, inherited from the journal article, will serve a lay audience poorly and will often be misread. Deliberate choices, made with an understanding of how audiences work and a firm commitment to the integrity of the evidence, are what this booklet sets out to teach. Chapter 1: Beyond the Deficit: What Audiences Actually Do with Science In 1985 the Royal Society published a report, chaired by the geneticist Walter Bodmer, called The Public Understanding of Science. It argued that a scientifically literate public was essential to a modern democracy and that scientists had a duty to communicate with it. The report was influential and well-intentioned, and it helped launch a generation of public lectures, science festivals and popular books. It also crystallised a set of assumptions that have proved remarkably durable: that the public's relationship with science is chiefly a matter of knowledge, that scepticism or opposition reflects a shortage of that knowledge, and that the remedy is for experts to supply more of it, more clearly. Scholars later gave this cluster of assumptions a name: the deficit model. Fifteen years after Bodmer, the House of Lords Select Committee on Science and Technology published a report, Science and Society (2000), written in the aftermath of the BSE crisis, in which British officials had spent years reassuring the public that beef was safe before the government announced in 1996 a probable link between BSE and a new variant of Creutzfeldt-Jakob disease in humans. The Lords described a crisis of confidence and called for a shift from one-way communication towards dialogue. The phrase "from deficit to dialogue" became a slogan of the field. And yet, as any researcher who has sat through a media-training session or read a university's public engagement strategy can confirm, the deficit model has never gone away. Understanding why it fails, and why it persists anyway, is the necessary starting point for everything else in this booklet. Why more facts rarely settle the matter The deficit model makes a testable prediction: people who know more science should hold attitudes closer to the scientific consensus. The evidence does not support it in any strong form. In a widely cited 2004 meta-analysis of survey data from Europe and the United States, Patrick Sturgis and Nick Allum found that scientific knowledge is associated with attitudes towards science, but modestly, and that the relationship varies considerably by topic and is shaped by political and cultural factors. Knowledge matters; it is simply one input among several, and often not the decisive one. The more striking result came from work on what Dan Kahan and colleagues call cultural cognition. In a 2012 paper in Nature Climate Change, Kahan's group reported that among Americans, higher scores on measures of science literacy and numeracy were not associated with greater concern about climate change overall. Instead, they were associated with greater polarisation: the most scientifically literate people on each side of the cultural divide were the furthest apart. The interpretation that has attracted the most attention is that people use their reasoning skills to defend positions that signal membership of groups they care about. If accepting a finding would put you at odds with your community, then being cleverer mainly makes you better at finding reasons to reject it. This literature is sometimes over-read. It does not show that facts are useless or that everyone is hopelessly tribal on every topic. Most scientific findings never become identity markers, and on those topics clear information does shift understanding. Later work has also debated how general the polarisation effect is. But on the minority of issues that have become entangled with political or religious identity, such as climate change, evolution, vaccination in some communities, nuclear power and genetically modified crops, the lesson is sobering. A communicator who treats resistance as ignorance will misdiagnose the problem, and a message built on that diagnosis may make things worse by implying that the audience is stupid. Why, then, does the deficit model survive? A 2016 paper by Molly Simis, Haley Madden, Michael Cacciatore and Sara Yeo in Public Understanding of Science examined this question directly. They suggested several reasons: scientists' training emphasises rational, evidence-based argument and may lead them to assume others reason the same way; the model is simple and flatters the expert's role; and institutional structures, including how communication is funded and evaluated, reward counting outputs such as lectures delivered and articles written rather than changes in understanding or relationships. In other words, the deficit model persists less because anyone defends it on the evidence than because it is the path of least resistance. The alternatives that have grown up since the Lords report are often grouped into three broad families, summarised in Table 1. The categories overlap in practice, and most real engagement mixes them, but it is useful to know which assumptions you are operating under. Table 1. Three broad models of public engagement with science. Model Assumption about the public Role of the researcher Typical formats Main limitation Deficit Lacks knowledge; scepticism signals ignorance Expert who transmits facts Lectures, press releases, explainers Ignores values, identity and trust Dialogue Has relevant knowledge, concerns and values Participant in two-way exchange Q&A events, consultations, cafés Can become consultation without influence Participation Has a legitimate stake in shaping research and its uses Co-producer and partner Citizen juries, co-design, citizen science Costly, slow, hard to scale The dialogue and participation models are not simply more virtuous versions of the deficit model. They answer different questions. If a research council wants to know how the public weighs the benefits and risks of a new technology before it is widely deployed, an approach that James Wilsdon and Rebecca Willis described in their 2004 Demos pamphlet See-through Science as "upstream engagement", a lecture explaining the technology will not produce that knowledge. If a hospital wants patients to understand how to take a new medication safely, a deliberative citizens' jury is overkill. The right model depends on the purpose, and the first question a communicator should ask is what the engagement is actually for. Dialogue has its own failure modes, and they are worth knowing because they can damage trust as badly as a patronising lecture. The United Kingdom's 2003 national public debate on genetically modified crops, run under the title GM Nation?, was an ambitious attempt to put dialogue into practice at scale, with regional meetings, local events and published findings. A detailed independent evaluation by Tom Horlick-Jones and colleagues, published in 2007 as The GM Debate: Risk, Politics and Public Engagement, found real strengths but also problems: the open meetings tended to attract people who already held strong views, the timetable was compressed, and it was unclear to participants how their input would affect decisions. The general lesson has been repeated in many settings since. Participants who suspect that a consultation is a ritual, with the decision already made, conclude that they were being managed rather than heard. If you invite people into dialogue, be clear at the outset about what is genuinely open to influence and what is not, and report back afterwards on what happened to their contributions. There is also a quieter form of deficit thinking that survives inside dialogue formats. A "café scientifique" or a public question-and-answer session can look like two-way exchange while functioning as a lecture with a longer queue at the end. The test is whether anything the audience says can change what the researcher thinks, does or says next time. If the honest answer is no, the event is still worth running, but it should be designed and described as what it is. Finding out what people already believe If the facts do not simply pour into empty vessels, what does happen to them? The most practically useful answer comes from risk communication research, and in particular from the "mental models" approach developed at Carnegie Mellon by M. Granger Morgan, Baruch Fischhoff, Ann Bostrom and Cynthia Atman and set out in their 2002 book Risk Communication: A Mental Models Approach. The premise is that people do not lack beliefs about technical subjects; they have beliefs, often coherent and reasonable on their own terms, that differ from expert understanding in specific and discoverable ways. Communication fails when it ignores those beliefs, because new information is interpreted through them. The method has a clear sequence: Build an expert model of the problem: a structured representation, often an influence diagram, of the processes and factors that matter for the decisions people face. Conduct open-ended interviews with members of the intended audience, beginning with very general prompts ("Tell me about radon") and only gradually narrowing, so that their own framings emerge rather than the interviewer's. Compare the lay and expert models to identify gaps, misconceptions and, importantly, correct beliefs that the communication can build upon. Use a structured questionnaire with a larger sample to estimate how common each belief is. Draft a communication that targets the consequential gaps, then test it with members of the audience and revise. In their early work on household radon, the Carnegie Mellon group found that many people conflated radon with other forms of radiation and contamination and believed, for example, that a house with a radon problem would be permanently contaminated. That belief matters for behaviour: if contamination is permanent, testing seems pointless and frightening, whereas in reality radon levels can be reduced by relatively simple ventilation measures. A generic leaflet explaining what radon is and where it comes from would have left that misconception untouched. A leaflet designed after the interviews could address it directly. Fischhoff later summarised the broader programme in a 2013 paper in the Proceedings of the National Academy of Sciences, "The sciences of science communication". He described four tasks: identify the science most relevant to the decisions people face; determine what people already know; design communications to fill the critical gaps; and evaluate whether the communication is adequate, then repeat. The emphasis on decisions is essential. Scientists tend to communicate what they find most interesting or most novel. Audiences need what bears on the choices in front of them. Those are often not the same thing, and the difference explains a great deal of wasted effort. Most researchers will not run a formal mental models study before every talk, but the logic scales down. Before preparing any significant piece of public communication, it is worth answering five questions in writing: • Who, specifically, is the audience? "The general public" is not an answer; "parents of primary-school children in a district considering a new school air-quality policy" is. • What decision, if any, do they face, and on what timescale? • What do they probably already believe about the topic, including things that are correct, partly correct and mistaken? How could you find out cheaply, for instance by reading comment threads, talking to a teacher, or asking five people? • What do they value, and which of those values does the topic touch? • What reasons might they have to trust or distrust you, your institution or your field? Suppose, as a hypothetical example, that a team of atmospheric chemists has found that levels of fine particulate matter outside several schools peak during drop-off and pick-up times, largely because of idling vehicles. Working through the five questions, they might realise that their primary audience is not "the public" but parents and head teachers; that the decision in question is whether to support a car-free zone around the school gates; that many parents probably believe outdoor air is healthier than indoor air and that their own car contributes little; that they value their children's health, convenience and not being judged; and that a university team telling them to change their behaviour may be received as condescending. Each answer reshapes the communication. The team might lead with the timing finding rather than with average pollution levels, explain why a single car contributes more at close range than parents expect, and involve a local head teacher as a co-presenter rather than arriving as outside experts. Trust is the medium, not the message The third lesson of the post-deficit literature is that the source matters as much as the content. People rarely evaluate scientific claims directly; they cannot check the data themselves and would not have time if they could. Instead, they make judgements about whether the person or institution offering the claim is trustworthy, and they accept or reject the claim largely on that basis. This is not irrational. It is how everyone, including scientists outside their own specialisms, handles the division of cognitive labour in a complex society. Research on what makes an expert seem trustworthy has converged on a small set of dimensions. Friederike Hendriks, Dorothe Kienhues and Rainer Bromme developed an instrument, published in 2015 in PLOS ONE, that measures perceived expertise, integrity and benevolence. The distinction is useful. Expertise is whether you know what you are talking about. Integrity is whether you are honest and follow the rules of your profession. Benevolence is whether you have the audience's interests at heart. Susan Fiske and Cydney Dupree, writing in PNAS in 2014, reported survey evidence that Americans tend to see scientists as highly competent but only moderately warm, and suggested that communicators could earn trust by showing that they share the audience's goals rather than by displaying more credentials. The practical implications are substantial. Researchers typically work hard to establish expertise, which audiences largely grant already, and neglect integrity and benevolence, which are often where the doubt lies. Integrity is signalled by acknowledging uncertainty, disclosing funding and conflicts of interest, admitting when earlier views turned out to be wrong, and refusing to overstate. Benevolence is signalled by showing that you understand why the question matters to the people asking it, by listening before explaining, and by not treating disagreement as stupidity. A simple way to apply this is to audit a draft talk, article or briefing against the three dimensions before delivering it. Take the hypothetical air-quality team again. Their draft opens with the team's credentials and funding, spends most of its length on measurement methods and ends with a recommendation. An audit would show that expertise is covered several times over, integrity is barely addressed (the draft never mentions what the measurements cannot show, such as whether the peaks are large enough to affect health over a school year), and benevolence is absent (nothing acknowledges that the school run is hard, that many parents have no realistic alternative to driving, or that the team is not there to assign blame). Adding two honest sentences about limitations and one about the practical difficulty of the school run costs almost nothing in time and changes how the whole message is received. There is a further point about trust that researchers often miss: it is not only earned or lost by individuals. When a university press office hypes a modest finding, when a funder's campaign oversells a field's promise, when one scientist makes a confident prediction that fails, the cost is borne partly by everyone else in the field. Public trust in science is a commons. The BSE episode that prompted the Lords report was damaging not because officials lacked facts but because they had projected a confidence the evidence did not warrant, and when that confidence was shown to be misplaced, the damage spread to institutions well beyond the ministry involved. Each chapter that follows returns to this theme in a different register, because every technique in the communicator's kit, from narrative to framing to plain language, can either build this commons or deplete it. The shift from deficit thinking is therefore not a matter of abandoning facts. Facts remain the substance of what scientists have to offer. It is a matter of recognising that facts are received by particular people, with particular beliefs, facing particular decisions, who extend or withhold trust for particular reasons. A communicator who takes that seriously will spend more time listening and less time broadcasting, will choose different starting points, and will measure success by whether the audience is better able to understand and decide, not by whether it agrees. The next step is to give those facts a shape that human minds are built to follow. Chapter 2: The Shape of a Story: Narrative Structure for Research Ask a researcher to describe a project to a friend and you will often hear something like this: "We're looking at soil microbes in grassland plots, and we sampled at four depths, and we sequenced the communities, and we also measured nitrogen, and we found some interesting differences between the grazed and ungrazed plots, and we're writing it up now." Every clause is true. None of it is memorable, and the listener has no idea why any of it matters. The problem is not vocabulary or length. It is structure. The description is a list, and human beings are poor at remembering lists and good at remembering stories. Randy Olson, a marine biologist who left a tenured position to become a filmmaker, has spent much of the past two decades trying to persuade scientists of this. His first book on the subject, Don't Be Such a Scientist (2009), was a polemic about the cultural gap between science and mass communication. His second, Houston, We Have a Narrative (2015), was more practical: an attempt to give scientists a simple tool for building narrative structure into their communication. That tool, the ABT template, is the most useful single device in this booklet, and this chapter begins with it before turning to the research on why narrative works and where it becomes dangerous. And, But, Therefore Olson credits the core idea to an unlikely source: the creators of the animated series South Park, Trey Parker and Matt Stone, who described in a widely circulated classroom appearance how they revise scripts by replacing the word "and" between story beats with "but" or "therefore". A sequence of events joined by "and then" is dull; a sequence joined by "but" and "therefore" creates tension and consequence. Olson generalised this into a three-part template: We know that [setup] AND [further setup], BUT [problem, contradiction or gap], THEREFORE [what we did or what follows]. The "and" section establishes the context and agreement: the things the audience accepts as true. The "but" introduces a problem, a contradiction, a surprise or a gap in knowledge. That word is where the story begins, because it creates a question in the listener's mind. The "therefore" resolves the tension, at least partially, by describing what was done, what was found or what needs to happen next. Apply it to the soil-microbe description: Grasslands store a large share of the world's soil carbon, and grazing by livestock is the most common use of grassland. But we know surprisingly little about how grazing changes the microbial communities that control whether that carbon stays in the ground. Therefore we compared microbes in grazed and ungrazed plots at four soil depths, and found that the biggest differences were not at the surface, where most studies sample, but deeper down. Every fact in the original is still present or implied. What has changed is that the listener now knows why the work exists and what question it answers, and the finding lands as a resolution rather than as one more item on a list. The template also exposes weaknesses: if you cannot write a convincing "but", you may not yet know what your project's question really is. Olson makes this point forcefully. The ABT is not only a presentation device; it is a diagnostic for clarity of thought. Olson places the ABT in the middle of what he calls a narrative spectrum. At one end is "AAA", "and, and, and": pure accumulation of facts with no tension. This is the default mode of much scientific writing and almost all poorly prepared talks. At the other end is what he calls "DHY", "despite, however, yet": a structure with so many reversals and qualifications that the listener loses the thread. Many discussion sections live here. The ABT sits in between, with one clear problem and one clear response. For a lay audience, one "but" is usually enough. Two or three can work in a long-form piece, provided each is resolved before the next is introduced. Olson also proposes a scalable exercise he calls the Word-Sentence-Paragraph method. You first try to reduce your project to a single word that captures its central theme (for the grassland study, perhaps "depth"). Then a single sentence, often in the "nothing in X makes sense except in the light of Y" form borrowed from Theodosius Dobzhansky's famous 1973 essay title "Nothing in Biology Makes Sense Except in the Light of Evolution". Then a paragraph in ABT form. The point of the progression is to force prioritisation. Researchers usually know far more about their project than they can communicate; the discipline lies in choosing the one thing that everything else serves. The ABT has limitations, and Olson himself is clear about them. It is a structure for the core of a message, not a complete communication plan. It does not tell you which audience to address, which frame to choose or which words to use. It can be applied mechanically, producing a string of fake tensions ("Cells are important, but we don't know everything about them"). A good "but" is specific and genuine: it names an actual gap, contradiction or surprise that the audience can feel. Compare two versions for a hypothetical study of sleep in teenagers. Weak: "Sleep is important for health, but there is still much we don't know about it." Strong: "Teenagers need more sleep than adults, and most secondary schools start earlier than primary schools, but no one had measured what happens to exam performance when a school moves its start time later." The first "but" could be attached to any topic in biology; the second creates a specific question that the audience immediately wants answered, and it sets up a "therefore" that can report the actual study design and result. A useful test is to read your "but" aloud to someone outside your field and ask whether they can guess what the "therefore" will be about. If they can, the tension is real. And the "therefore" must be honest. If the study did not resolve the problem, the "therefore" should say what it did achieve and what remains open. There is also some evidence that narrative structure matters even inside the academy. In a 2016 study in PLOS ONE, Ann Hillier, Ryan Kelly and Terrie Klinger analysed the abstracts of several hundred climate-change papers and rated them for narrative features such as a sense of setting, a clear conflict or problem, and a resolution. Papers whose abstracts had more narrative characteristics tended to be cited more often, even after accounting for other factors they could measure. The study is correlational and cannot show that narrative writing causes more citations, but it undermines the assumption that narrative is only for outsiders. Why narrative works Michael Dahlstrom's 2014 review in PNAS, "Using narratives and storytelling to communicate science with nonexpert audiences", is the best short introduction to the research on why stories are so effective. He draws on a distinction made by the psychologist Jerome Bruner in the 1980s between two modes of thought. The paradigmatic or logical-scientific mode seeks general truths through argument, evidence and abstraction. The narrative mode deals in particular experiences, arranged in time, with characters whose actions have causes and consequences. Science is built in the first mode. Most people, most of the time, think in the second. Dahlstrom summarises several consistent findings. Narratives tend to be easier to understand than expository text for non-expert readers, and they are recalled better. They generate more engagement and interest. And they can be more persuasive, partly because of a phenomenon that Melanie Green and Timothy Brock called "transportation" in a 2000 paper in the Journal of Personality and Social Psychology. When people are absorbed into a story, they become less likely to generate counterarguments and more likely to adopt beliefs consistent with the story, even when they know it is fiction. That last property explains both the power and the danger of narrative. On the positive side, a well-built story can carry an audience through material they would never read as an argument. On the negative side, a story can persuade independently of its evidential weight. A single vivid case can outweigh statistical evidence covering thousands. A striking illustration comes from the psychology of vaccine decisions. In a 2011 experiment published in Medical Decision Making, Cornelia Betsch and colleagues exposed participants to a simulated online forum about vaccination in which the base-rate statistics on adverse events were held constant but the proportion of personal stories describing harm was varied. Participants who read a higher proportion of such stories perceived vaccination as riskier and reported lower intention to vaccinate, even though the statistical information they saw was the same. The stories did work that the numbers did not, and in this case the work was misleading. The related "identifiable victim effect" is well documented. In a 2007 study, Deborah Small, George Loewenstein and Paul Slovic found that people donated more to help a single identified child than in response to statistics about millions in need, and that adding the statistics to the individual story actually reduced donations. Scientists often complain that the public responds to anecdotes rather than data. This research suggests that the complaint describes a feature of human cognition rather than a failing of particular audiences, and that communicators who refuse to use stories are ceding the ground to those who will. The ethics of telling science as a story If narrative is this powerful, a researcher who uses it takes on responsibilities. Dahlstrom and Shirley Ho addressed this directly in a 2012 paper in Science Communication, "Ethical considerations of using narrative to communicate science". They framed the issue around three questions that every communicator should answer before building a story. The first is the purpose of the communication. Is the aim to help people understand something, or to persuade them to believe or do something? Both can be legitimate. A public health agency trying to increase vaccination rates is engaged in persuasion; a researcher explaining how a vaccine was tested is engaged in explanation. But the ethical standards differ. Persuasion by narrative, which works in part by reducing counterarguing, deserves particular scrutiny, and it is more easily justified when the evidence is strong and the behaviour in question benefits the audience than when the evidence is contested or the benefit accrues to someone else. The second is accuracy, and here narrative raises specific problems. A story has a representativeness problem: the individual case at its centre may or may not be typical. A patient whose tumour vanished on a new drug makes a compelling story, but if the response rate in the trial was modest, the story conveys a false impression unless it is explicitly placed in context. Stories also impose causal clarity on events that may not have been so clear. Narrative's natural grammar is "this happened, therefore that happened", and it will tend to make correlations feel like causes. The communicator's job is to make sure that the "therefore" in the story is one the evidence supports. The third is whether narrative should be used at all. Dahlstrom and Ho note that some scientists regard narrative as inherently at odds with scientific norms, because it privileges the particular over the general. Their answer, and the one this booklet adopts, is that avoiding narrative does not produce neutral communication. It simply produces communication that audiences are less likely to understand or remember, leaving the field to others who may have fewer scruples. The better response is to use narrative with discipline. In practice, that discipline comes down to a short set of rules: • Make the story typical, or say that it is not. If you use an individual case, choose one that represents the central tendency of the evidence, or explicitly state how it differs. • Pair the story with the numbers. The story gives the audience a reason to care; the numbers tell them how much to care. Research on the identifiable victim effect suggests the combination may reduce emotional impact, but in science communication, accuracy is the constraint and impact is the objective, not the other way round. • Do not invent. Composite characters and reconstructed scenes are common in journalism and documentary, but they must be labelled. A researcher who presents a composite patient as a real one, or a hypothetical scenario as a real event, has crossed a line that damages trust when discovered. • Let the "therefore" match the evidence. If the study showed an association, the story's resolution should not imply that one thing caused the other. If the finding needs replication, the story should end with an open question, not a triumph. • Watch the villain. Stories need conflict, and it is tempting to supply it with an antagonist: an industry, a regulator, a rival theory. Sometimes that is accurate. Often it is not, and it can turn an explanatory message into a partisan one, with consequences taken up in the next chapter. Stories at different scales The ABT works at the level of a paragraph or a thirty-second pitch. Longer formats need more architecture. Three structures are useful to know. The first is the quest, which follows a researcher or team pursuing a question through obstacles to an answer. It suits profiles, documentaries and long-form features, and it is the structure most people imagine when they think of "science stories". Its strength is that it makes the process of science visible, including dead ends and uncertainty. Its weakness is that it foregrounds the scientist, which can feel self-promotional and can distort the collective nature of most research. The second is the mystery, which begins with a puzzling observation and works towards an explanation. This is often the most natural fit for scientific material, because much research really does begin with something that does not make sense. A talk that opens, in a hypothetical example, "Three summers ago, fishermen along one stretch of coast started pulling up species that should not have been there" creates immediate engagement, and the rest of the talk can unfold as the solution of the puzzle. The mystery structure has the virtue of mirroring how science actually works: observation, hypothesis, test, revision. The third is the problem and response, which begins with something that matters to the audience, explains its causes, and ends with what can be done. This is the natural structure for policy communication and public health messages, and it is essentially an extended ABT. Its risk is that the "response" may carry more certainty than the evidence warrants, especially if the communicator is also an advocate for a particular solution. Within any of these structures, a few narrative elements consistently repay attention. Characters give the audience someone to follow, and they need not be human: an organism, a molecule, a glacier or a dataset can serve if it is given a trajectory. Specific details, a place, a date, an object, anchor abstractions. Stakes tell the audience why the outcome matters. And a clear ending, even if it is an open question, gives the audience something to carry away. Consider how these elements might transform a hypothetical five-minute talk on antibiotic resistance for a school audience. The expository version would define antibiotics, explain the mechanism of resistance, present surveillance data and conclude with recommendations. A narrative version might follow a single bacterium through a course of antibiotics that the patient stops taking early, showing how the survivors are the ones that happen to carry resistance genes, and how they pass those genes on. The mechanism is the same; the audience now has a character to follow and a moment of consequence to remember. The surveillance data can then appear as the answer to a question the story has already raised: how often does this happen? Narrative, in short, is not an alternative to evidence. It is the vehicle that carries evidence into minds that would otherwise let it slide past. The communicator's craft lies in building a vehicle that holds attention without distorting the load. The next question is which part of the load to put at the front, and that is the problem of framing. Chapter 3: Framing: Choosing the Lens Without Distorting the Picture In 1981 Amos Tversky and Daniel Kahneman published a short paper in Science that has become one of the most cited in the social sciences. They asked participants to imagine that the United States was preparing for the outbreak of an unusual disease expected to kill 600 people, and to choose between two programmes. In one version of the question, participants were told that Programme A would save 200 people, while Programme B had a one-third probability of saving all 600 and a two-thirds probability of saving no one. Most chose A. In another version, the same options were described in terms of deaths: under Programme C, 400 people would die; under Programme D, there was a one-third probability that nobody would die and a two-thirds probability that 600 would die. Most chose D. The options were mathematically identical. Only the description had changed, from lives saved to lives lost, and preferences reversed. This is framing in its narrowest sense, sometimes called an equivalence frame: logically equivalent descriptions that produce different judgements. Scientists who communicate about risk encounter it constantly. A treatment that gives a 90 per cent survival rate sounds better than one with a 10 per cent mortality rate, although they are the same treatment. Health communicators have spent decades trying to turn the equivalence effect into practical guidance. In an influential 1997 paper in Psychological Bulletin, Alexander Rothman and Peter Salovey proposed that gain-framed messages ("exercising keeps your heart strong") should work better for encouraging prevention behaviours, while loss-framed messages ("failing to get screened can mean a cancer is found late") should work better for detection behaviours such as screening, because detection involves accepting the risk of an unwelcome finding. The idea was elegant, and it shaped a great deal of public health messaging. Subsequent meta-analyses by Daniel O'Keefe and Jakob Jensen, however, found that the differences in persuasiveness between gain and loss frames were generally small and not consistent across behaviours. The honest lesson for researchers is modest: the way a number is described matters, audiences should ideally be given both forms (survival and mortality, benefit and harm), and no single wording trick should be relied on to do the work of a clear explanation. The ethical point is sharper than the persuasive one. Because equivalence frames can shift choices without changing any facts, a communicator who presents only the favourable form of a statistic, survival without mortality or relative risk reduction without absolute risk, is steering the audience rather than informing it. Chapter 4 returns to this in detail. But most of what communication researchers mean by framing is broader, and more consequential, than this. What a frame does The political scientist Robert Entman offered a definition in 1993 that is still widely used: "To frame is to select some aspects of a perceived reality and make them more salient in a communicating text, in such a way as to promote a particular problem definition, causal interpretation, moral evaluation, and/or treatment recommendation." These are sometimes called emphasis frames. They do not describe the same thing in equivalent ways; they choose which of many true things about an issue to put in the foreground. Every scientific finding of public interest has multiple aspects. Consider research on gene drives, a genetic technique that can spread a trait through a wild population faster than normal inheritance would allow. It can be described as a potential tool for eliminating malaria-carrying mosquitoes, as an ecological intervention with uncertain downstream effects, as a question of who gets to decide whether a technology is released in a particular country, as a frontier of scientific competition, or as an example of humans altering nature in a new way. All of these descriptions are accurate. Each emphasises a different aspect, invites a different interpretation of what the issue is about, and appeals to different values. The communicator cannot avoid choosing among them, because no message can say everything. Even a deliberately neutral description, with no evaluative language, embeds a frame through what it mentions first and what it leaves out. The idea that scientists should think consciously about framing entered wide debate through a 2007 Science article by Matthew Nisbet and Chris Mooney, "Framing Science". They argued that scientists often assume the facts speak for themselves, and that this assumption leaves the framing of scientific issues to others, especially to those with strong interests in the outcome. Their article provoked a lively response. Some scientists wrote that framing sounded like spin, and that scientists should simply present the evidence accurately and let the public decide. Nisbet and Mooney's reply, and the position that most communication scholars now take, is that this objection misunderstands the choice available. Presenting "just the facts" involves selecting which facts, in what order, with what emphasis. The question is not whether to frame but whether to frame deliberately and honestly. Drawing on earlier work by William Gamson and Andre Modigliani and by John Durant and colleagues, Nisbet identified a set of frames that recur across science-related policy debates in many countries and topics. Table 2 presents a version of this typology, with a hypothetical illustration for each frame from the gene-drive example. Table 2. Recurring frames in public debates about science, with gene drives as an illustration. Frame Defines the issue as Hypothetical gene-drive illustration Social progress Improving quality of life or solving problems A tool that could help eliminate malaria Economic development Investment, competitiveness, jobs A field where national research capacity is at stake Morality and ethics Right and wrong, limits we should respect Whether humans should permanently alter wild species Scientific uncertainty What is known and not known Unknown ecological effects of removing a species Pandora's box Technology out of control, irreversible harm A release that cannot be recalled Public accountability Who decides and in whose interest Whether affected communities consented to field trials Middle way Compromise between polarised options Phased, contained trials with community oversight Conflict and strategy A game among elites; who is winning Rival research groups and funders competing Several observations follow from the table. First, the frames are not mutually exclusive, and serious communication usually combines several. Second, the frames carry different implications for what should be done: a social progress frame implies urgency, a Pandora's box frame implies caution, a public accountability frame implies consultation. Third, and most important for researchers, the frames resonate differently with different audiences. A community in a malaria-endemic region, a European environmental group, a national funding agency and an ethics committee will each find some frames more relevant than others. Evidence that frames change who listens The strongest evidence for the value of deliberate framing comes from studies showing that a change of frame can shift how particular audiences respond to the same underlying science. Climate change has been the main testing ground. In a 2010 study published in BMC Public Health, Edward Maibach, Matthew Nisbet and colleagues tested how different segments of the American public responded to a description of climate change framed as a public health issue, compared with environmental and national security framings. They found that the public health frame tended to elicit more hopeful and engaged reactions across segments, including among some audiences that were dismissive of climate change as an environmental problem. The study was exploratory, and later research has produced mixed results about how large and durable such effects are, but it helped establish public health as one of the standard frames in climate communication. A different line of research has examined how frames interact with moral values. Matthew Feinberg and Robb Willer, in a 2013 paper in Psychological Science, drew on moral foundations theory, which proposes that liberals and conservatives in the United States tend to weight different moral intuitions. Environmental messages typically emphasise harm and care, foundations that liberals prioritise. Feinberg and Willer found that reframing pro-environmental messages in terms of purity and sanctity, emphasising pollution as contamination of something pure, increased pro-environmental attitudes among conservatives and substantially narrowed the gap between the two groups. The finding suggests that part of what looks like resistance to science is a mismatch between the values the message invokes and the values the audience holds. Dan Kahan and colleagues reported a related effect in a 2015 study in The Annals of the American Academy of Political and Social Science. They found that reading about geoengineering, the deliberate large-scale intervention in the climate system, as part of a story about climate science made some participants who were otherwise sceptical more open to the evidence on climate change itself. Their interpretation was that emphasising technological responses, which are compatible with the values of people who favour markets and innovation, reduced the identity threat that climate evidence ordinarily carried for them. Whatever one thinks of geoengineering as policy, the study illustrates the principle: the frame around a fact can determine whether an audience engages with the fact at all. The messenger is part of the frame. A frame that emphasises religious stewardship of creation will be received differently if it comes from a researcher who shares the audience's faith than from one who does not. The climate scientist Katharine Hayhoe, an evangelical Christian, has for many years spoken to faith communities about climate change in terms of shared values, and her work is often cited as an example of the importance of messenger-audience fit. For most researchers, the practical lesson is not to adopt identities they do not hold, which would be dishonest and quickly detected, but to recognise when a trusted intermediary, a local doctor, a farmer, a faith leader or a teacher, would be a more credible messenger than a university scientist, and to support those intermediaries rather than competing with them. Where framing becomes distortion If framing is unavoidable and can make communication more effective, where are its limits? This is not an abstract concern. Framing techniques are used every day by lobbyists, campaigners and political strategists whose goal is to win, not to inform, and the same techniques can be turned towards conclusions the evidence does not support. A workable line can be drawn with three tests. The first is accuracy of emphasis. An emphasis frame selects among true aspects of an issue. It becomes distortion when it implies something false about the overall balance of evidence. Framing gene drives primarily around malaria elimination is legitimate if the communicator is honest about how far the technology is from deployment and what uncertainties remain. It becomes misleading if it implies that elimination is imminent or assured. The frame may highlight; it may not misrepresent. The second is consistency across audiences. It is legitimate to emphasise the health aspects of air pollution to a health audience and the economic aspects to a business audience. It is not legitimate to tell one audience that a risk is serious and another that it is negligible. A useful check is to ask: if these two audiences compared notes, would they feel deceived? If the answer is yes, the framing has crossed a line. The third is respect for the audience's decision. Framing to help people understand how an issue connects to their own values is a service. Framing to exploit a cognitive bias so that people reach a conclusion they would reject on reflection is manipulation. The difference is often visible in intent: is the communicator trying to make the evidence meaningful, or trying to prevent the audience from thinking about it carefully? Researchers face one framing temptation more than any other: the social progress frame applied to their own work. Grant applications, press releases and institutional communications all reward describing research as a step towards a cure, a solution or a breakthrough. Each individual instance may be defensible; the cumulative effect is a public that hears about cures that never arrive and begins to discount everything scientists say. Chapter 6 examines research on how exaggeration in press releases passes into news coverage. The framing lesson is that the social progress frame is the one most likely to generate cynicism when overused, and researchers should reach for it with particular care. There is also a risk in adopting the conflict frame, which journalists find attractive because conflict is newsworthy. A researcher who is drawn into describing a debate as a battle between their side and another, whether a rival scientific camp, an industry or a political party, may gain attention in the short term but will tend to push the issue towards polarisation, where, as Chapter 1 noted, facts do less work. When a journalist asks "Who is winning this debate?", a better response usually redirects to what the evidence shows and what remains uncertain. Choosing a frame in practice Suppose, as a hypothetical example, that a research team has completed a study showing that a common agricultural fungicide reduces the foraging efficiency of wild bumblebees at concentrations found in field margins. The team is invited to present the findings at three events in the same month: a regional farmers' association meeting, a public lecture at a natural history museum, and a briefing for officials at the national agricultural ministry. A deficit-model approach would prepare one talk and deliver it three times. A framing-aware approach begins with the audience analysis described in Chapter 1 and then asks which true aspects of the finding are most relevant to each. For the farmers, the natural frame is economic development combined with middle way. Many crops depend on pollination, and farmers have direct economic interests in healthy pollinator populations. The team might lead with the dependence of certain local crops on wild bumblebees, present the finding as a risk to a service farmers rely on, and discuss practical options, such as timing of application or buffer strips, that reduce exposure without abandoning the fungicide. Leading with a Pandora's box frame ("this chemical is harming nature") would be accurate in part but would likely trigger defensiveness and lose the room. For the museum audience, social progress in the sense of understanding, and scientific uncertainty, may be more engaging: how the team measured foraging, what surprised them, what they still do not know about how the effect scales to whole populations. This audience wants a story of discovery, and the mystery structure from the previous chapter fits well. For the ministry, public accountability and scientific uncertainty are central. Officials need to know how strong the evidence is, how it relates to the existing regulatory assessment, what further evidence would change the picture, and what the range of policy options looks like. As Chapter 5 will discuss, this is where the researcher also has to decide what role to play. How would the team know whether its framing choices were right? Large organisations test frames with surveys and focus groups; a research team rarely has that luxury, but it can do a scaled-down version. A draft of the farmers' talk could be shown to one or two agronomists or extension officers who work with farmers daily, with a simple request: what would make this audience stop listening? The museum version could be tried on a few friends outside science. The ministry briefing could be read by a former civil servant or a colleague who has sat on an advisory committee. None of this is rigorous research, but it catches the most common failure, which is a frame that seems natural to the researcher and alienating to the audience. It also often reveals frames the team had not considered: an agronomist might point out, for instance, that local farmers are already worried about declining yields in a particular crop and would be more receptive if the talk connected the finding to that concern. In all three settings, the core findings are identical, the limitations are stated, and none of the audiences would feel deceived if they compared notes. What differs is the entry point and the emphasis. That is framing done responsibly: not changing the evidence to suit the audience, but finding the door through which each audience can walk into it. Hashtags: #ScientificStorytelling #PublicEngagement #ScienceCommunication #PublicUnderstandingOfScience #DeficitModel #DialogueModel #ParticipatoryEngagement #AudienceAnalysis #MentalModels #TrustInScience #RiskCommunication #NarrativeCommunication #ABTStorytelling #ScienceFraming #EvidenceBasedCommunication #CommunicatingUncertainty #SciencePolicyCommunication #MediaEngagement #PressReleases #VideoAbstracts #SocialMediaScience #MisinformationResponse #PublicTrust #ResearchCommunication #FutureOfScienceEngagement

  • Scientific Crowdfunding (Launching Campaigns, Community Backing, and Public Science)

    Download the Book (PDF): Introduction A graduate student in ecology wants to know whether a population of salamanders in a single mountain stream carries a fungal pathogen that has devastated amphibians elsewhere. The question is sharp, the fieldwork would take a month, and the laboratory work would cost a few thousand dollars in swabs, reagents and sequencing. No federal agency will fund it. The study is too small to be worth a panel's time, too preliminary to have the pilot data a panel expects, and too uncertain to guarantee a publishable result. Her adviser's grant covers other work. Her department's travel fund covers conferences, not field seasons. The question will go unasked unless she finds the money somewhere else. Over the past fifteen years, a growing number of researchers in that position have found it in the same place: among strangers, friends, former teachers, amateur naturalists, patients, retirees and the idly curious, who give ten, fifty or five hundred dollars through a crowdfunding page because they want the question answered. Dedicated platforms now exist for exactly this purpose. The largest of them, Experiment, reports that roughly half the campaigns it hosts reach their targets, and the typical target is a few thousand dollars. Larger efforts have raised far more. The American Gut Project, a citizen-science study of the human microbiome run from the University of California, San Diego, drew in more than a million dollars from participants who paid to have their own samples sequenced and contributed to a shared, open dataset. This book is about how to do that well, and about why doing it well is more demanding than it first appears. The argument of this book The usual way of describing scientific crowdfunding treats it as a small grant from an unusual source. You write a short proposal, record a video, set a target, and if enough people give you money, you have a grant. The difference from an agency grant, on this view, is mainly one of scale and reviewer: a crowd of laypeople instead of a panel of experts. That description is wrong in a way that causes real harm. A crowdfunded research project is not a grant. It is a public promise made to a community that the researcher has, in large part, assembled. The money arrives attached to people, and the people arrive with expectations: that they will be told what happens, that the work will actually be done, that the results will be shared whether or not they are exciting, and that the researcher is the kind of person they believed when they clicked the button. None of those expectations is written into a contract. All of them are real. A researcher who treats the campaign as a funding transaction and then disappears into the laboratory will, sooner or later, discover that the community she built can turn against her, and that the damage extends to every scientist who tries the model after her. The controlling idea of this book follows from that. Crowdfunding works for science, and remains honest, when the whole project is designed around the relationship with backers from the beginning: the choice of question, the budget, the pitch, the timeline and the eventual publication all have to be built for an audience that is paying to watch. Researchers who design this way find that the discipline improves their science. They scope questions more tightly, budget more carefully, state their risks more plainly and publish their results more openly than they might under a conventional grant. Researchers who do not design this way tend to fail twice: first at raising the money, and then, if they succeed, at keeping faith with the people who gave it. What crowdfunding is good for It is worth being clear from the start about the kind of research this model serves. Crowdfunding is well suited to small, well-bounded, high-risk exploratory questions: a pilot study, a field season, a sequencing run, a prototype instrument, a survey, a first attempt at a method nobody has tried. It suits early-career researchers without track records, independent researchers without institutions, researchers in unfashionable fields, and researchers whose questions cross disciplinary lines that grant panels police. It suits work with a story that ordinary people can follow and care about. It is poorly suited to large, capital-intensive projects, to anything that requires years of guaranteed salary, to clinical trials of treatments, and to work whose value can be explained only to specialists. It is dangerous when it is used to promise outcomes that research cannot promise, such as cures, products or guaranteed discoveries. Much of what follows is about recognizing which side of that line a given idea sits on, and reshaping ideas so that they land on the right side. How the book is organized The chapters follow the life of a campaign, from the decision to try it to the moment the last backer receives the final report. Chapter 1 explains why crowdfunding for research exists at all: the structural reasons conventional funding systems struggle to support small exploratory work, the short history of research-specific platforms, and what the published evidence says about which campaigns succeed. Chapter 2 is about choosing and shaping the project itself, which is where most campaigns are won or lost long before launch. Chapter 3 surveys the platforms, funding models and rules, including the institutional and legal questions that a researcher inside a university and an independent researcher will answer differently. Chapter 4 is about money: building a campaign budget that accounts for fees, rewards, taxes and contingency, and setting a target that is both honest and achievable. Chapter 5 is about the audience, which the evidence suggests matters more than anything else, and about the weeks of preparation that precede a good launch. Chapter 6 covers the pitch video and the campaign page, the two objects most prospective backers will actually see. Chapter 7 turns to what happens after the money arrives: updates, delays, failures, negative results and the management of expectations over a project that may take longer than anyone planned. Chapter 8 addresses the ethical duties specific to research funded by the public, including the oversight of work with human participants or animals, the handling of data, the honest reporting of results and the obligations of researchers who work outside institutions. The conclusion draws out what all of this implies for researchers, for institutions and for the wider relationship between science and the public that pays for it. Throughout, the book draws on the small but real body of research on scientific crowdfunding, on well-documented public cases, and on the practical experience that has accumulated since the first research campaigns in the early 2010s. Where the evidence is thin, it says so. Where a practice is a matter of judgment rather than data, it says that too. A word on scale Nothing in this book assumes that crowdfunding will replace public research funding, and the argument does not depend on it doing so. The sums involved are small. A successful campaign typically covers a few thousand dollars of direct costs, not salaries, not overheads, not a laboratory. The point is not that the crowd can fund science in general. The point is that there is a specific, valuable kind of research that the established system funds badly, and a specific, well-run kind of campaign that can fund it well, while opening the work to people who would otherwise never see how research is done. That second benefit is not a side effect. For many researchers who have run campaigns, the most lasting result was not the money but the audience: a few hundred people who now follow the work, ask questions, forward papers to their friends and, sometimes, give again. Treated with care, that audience is one of the most valuable things a researcher can build. Treated carelessly, it becomes a liability. The chapters that follow are about the difference. Chapter 1: Why the Crowd Funds What Panels Will Not Every funding system has a shape, and the shape determines which questions get asked. Before deciding whether to crowdfund a piece of research, it helps to understand why the established system leaves certain kinds of work unfunded, because the gaps it leaves are precisely where crowdfunding does its best work. A researcher who understands the gap can design a campaign to fill it. A researcher who does not will often try to use crowdfunding as a cheaper, easier substitute for a grant, and will be disappointed on both counts. The shape of conventional research funding Most research in wealthy countries is funded through competitive grants awarded by government agencies, foundations and, increasingly, philanthropic organizations. The core mechanism is peer review: a proposal is read by a small number of experts, scored, discussed by a panel, and ranked against other proposals. Money flows to the highest-ranked proposals until the budget runs out. This system has real strengths. It concentrates judgment in people who understand the field, it provides a check against obvious error, and it spreads money across many institutions rather than letting it pool around the politically connected. But it also has predictable biases, and several of them work against small exploratory research. The first is conservatism. Reviewers are asked to judge whether a project will succeed, and the most reliable evidence of future success is past success: preliminary data, a track record in the method, a publication history in the area. A genuinely new idea, by definition, lacks these things. The economist Paula Stephan, in her book How Economics Shapes Science, describes at length how the incentives of grant competition push both applicants and reviewers toward projects that are safe enough to promise results. When funding rates are low, the pressure intensifies. At the United States National Institutes of Health, success rates for the main investigator-initiated research grants have hovered around one application in five in recent years, and reviewers faced with that arithmetic have little room to back a long shot. The second bias is scale. The fixed costs of writing, reviewing and administering a grant are substantial, and they do not shrink much when the grant does. It is uneconomic for an agency to run a full review process for a request of three thousand dollars, and so most agencies do not offer awards that small, or offer them only through narrow schemes. The researcher who needs a few thousand dollars for a field season or a sequencing run falls below the floor of the system. The third bias is eligibility. Most major funders restrict applications to people with particular appointments at particular kinds of institutions. Graduate students usually cannot apply as principal investigators. Independent researchers, community scientists and people between positions usually cannot apply at all. Researchers at small teaching institutions can apply but compete against research universities with dedicated grant offices. The fourth is time. From idea to money, a conventional grant commonly takes the better part of a year: writing, submission deadlines, review cycles, council decisions, award paperwork. For a question tied to a season, an event, an outbreak or a short-lived opportunity, that delay can be fatal. None of this makes grant funding bad. It makes it a system optimized for a particular kind of project: medium to large, built on established foundations, led by credentialed investigators, and able to wait. The kind of project this book is about is the complement of that: small, new, sometimes led by people outside the usual credentials, and often in a hurry. Where crowdfunding fits Crowdfunding inverts several features of the grant system. The reviewers are not experts but interested members of the public, and they are not asked to rank proposals against one another but to decide, individually, whether a particular project is worth some of their own money. There is no formal eligibility requirement beyond what the platform imposes. Campaigns can be prepared in weeks and run for a month. And because each backer gives a small amount, the total can be as small as the project requires. These features make crowdfunding naturally suited to the gaps just described. A graduate student can run a campaign in her own name. An independent researcher can raise money without an institutional affiliation. A project that needs three thousand dollars can ask for three thousand dollars without apology. A question tied to a season can be funded before the season arrives. But the inversion also creates new constraints. Backers are not experts, and they cannot evaluate methodology the way a panel can. What they can evaluate is whether the question interests them, whether the researcher seems competent and honest, and whether the plan seems concrete enough to be carried out. This means that a crowdfunded project must be legible: the question must be explainable to a curious non-specialist, and the connection between the money and the answer must be visible. Research that can only be justified in the technical language of a field, however important, is hard to crowdfund. The absence of expert review also shifts responsibility. When a panel funds a project, the panel shares in the judgment that the project is sound. When a crowd funds a project, the judgment rests almost entirely on the researcher's own representations. This is one reason the ethical demands on crowdfunding researchers are, in some respects, higher than on grant recipients, a theme Chapter 8 develops at length. A short history of research crowdfunding Crowdfunding as a general practice predates the internet: subscription publishing, community barn-raisings and public appeals for monuments all pooled small contributions for a shared purpose. The modern online form took off around 2009, when Kickstarter launched, joining earlier platforms such as Indiegogo. These general-purpose platforms were designed for creative projects and products, and their model was built around rewards: backers gave money and, in return, received the album, the gadget, the book or some token of thanks. Scientists noticed quickly. Some early research campaigns ran on the general platforms, and a handful attracted wide attention. In 2012, the pharmacologist Ethan Perlstein, then a postdoctoral researcher at Princeton, raised roughly twenty-five thousand dollars on the platform RocketHub for a study of how methamphetamine distributes inside neurons, a project that drew media coverage partly because of its provocative subject and partly because it was among the first to show that a basic-science question could attract public money. Perlstein went on to found an independent research company and remained a prominent advocate of alternative funding for science. At around the same time, researchers began organizing collective experiments in science crowdfunding. The #SciFund Challenge, started in 2011 by the ecologists Jarrett Byrnes and Jai Ranganathan, ran a series of coordinated rounds in which groups of scientists launched campaigns at the same time, shared advice and compared results. The organizers treated the rounds as a natural experiment and later published an analysis of what distinguished the campaigns that raised more money from those that raised less. Their central finding, discussed in detail in Chapter 5, was that the size of a researcher's existing audience mattered more than almost anything about the campaign itself. Dedicated research platforms followed. Petridish, launched in 2012, hosted science projects for a few years before closing. Microryza, founded in 2012 by Cindy Wu and Denny Luan, rebranded as Experiment and became the largest platform specifically for research. Other platforms emerged for particular communities: academist in Japan, for example, and Consano, which focused on medical research. Universities began running their own crowdfunding sites, often through their development offices, to let researchers raise money from alumni and local supporters. The same period produced a few larger and more complicated cases. The American Gut Project, launched in 2012 by the microbiologist Rob Knight and colleagues, invited members of the public to contribute money in exchange for having their own microbiome sequenced, and assembled what became one of the largest open human microbiome datasets; its organizers published an overview of the project in the journal mSystems in 2018. Around the same time, a start-up called uBiome raised money on Indiegogo with a similar citizen-science pitch; it later grew into a venture-backed company whose founders were charged by federal authorities in 2021 with securities and health-care fraud over its later billing of insurers and its dealings with investors. The contrast between the two is instructive and is revisited in Chapter 8. And in 2013 the Glowing Plant project raised more than four hundred and eighty thousand dollars on Kickstarter by promising backers seeds of a genetically engineered plant that would glow; the plants never arrived, and the episode led Kickstarter to prohibit genetically modified organisms as rewards. Chapter 7 examines what went wrong. What the evidence says about success Research on scientific crowdfunding is still modest in volume, but a few careful studies give a reasonable picture of what a typical campaign looks like and what distinguishes successful ones. The most comprehensive is an analysis by Henry Sauermann, Chiara Franzoni and Kourosh Shafi, published in PLOS ONE in 2019, of 725 campaigns on Experiment. About 48 percent of campaigns reached their targets. The median target was around three thousand five hundred dollars, and the median amount raised by successful campaigns was a little over three thousand. Around four-fifths of campaign creators were affiliated with educational institutions, and many were students: undergraduates, master's students and doctoral students together made up more than half. Most projects were straightforward research, concentrated in biology, ecology, medicine and engineering. Several features were associated with success. Campaigns with smaller targets were more likely to be funded, as one would expect. Campaigns that posted lab notes, the platform's name for progress updates, and campaigns that included a video, did better. Strikingly, students and postdoctoral researchers had higher success rates than professors, and campaigns led by women had higher success rates than campaigns led by men. The authors were careful not to claim that these associations are causal; many unobserved factors, such as how much effort a creator put into outreach, could explain them. But the pattern is suggestive. The attributes that drive grant success, such as seniority and publication record, did not drive crowdfunding success. Effort, communication and connection to an audience appeared to. This fits the earlier #SciFund analysis. Byrnes, Ranganathan and their colleagues, writing in PLOS ONE in 2014, found that the amount a campaign raised tracked the reach of the researcher's outreach, for example the number of people who visited the campaign page and the size of the researcher's existing online following. A researcher who had spent years talking to the public about her work had, in effect, already done most of the fundraising before the campaign opened. A researcher who had not was starting from nothing. A useful comparison comes from studies of crowdfunding in general. Ethan Mollick, whose 2014 study in the Journal of Business Venturing was one of the first systematic analyses of Kickstarter, found that project quality signals and the size of the founder's social network both predicted success, and that most funded projects delivered what they promised, though often late. A later study of delivery by Mollick, based on a large survey of Kickstarter backers and released in 2015, estimated that around nine percent of successfully funded projects failed to deliver their rewards at all. Research projects are different from product launches in many ways, but the lesson carries over: funding is only the beginning, and delay is the norm rather than the exception. The real currency Put together, the history and the evidence point to a conclusion that shapes the rest of this book. The currency of scientific crowdfunding is not the quality of the proposal as an expert would judge it. It is trust, and trust is built from three things: an audience that already knows the researcher, a project that ordinary people can understand and care about, and a record of honest communication. That conclusion has an uncomfortable implication. Crowdfunding does not level the playing field as completely as its early advocates hoped. It removes some barriers, such as institutional eligibility and seniority, but it introduces others, such as the need for an audience and the ability to communicate. A brilliant researcher who has never spoken to the public, works on a question that cannot be explained in a paragraph and has no network of supporters will struggle. The advantages that crowdfunding confers flow to those who can build a relationship with a community. The implication is also an opportunity. Unlike seniority or institutional prestige, the ability to build an audience can be learned and practiced, and it compounds. A researcher who runs one honest, well-communicated campaign emerges with more backers, more followers and more credibility than she started with, and her second campaign begins from a stronger position. Several researchers who ran early campaigns have described exactly this pattern. The money from the first campaign was modest. The community it built was worth far more. When to consider crowdfunding Before going further, it is worth setting out the conditions under which crowdfunding is a sensible choice. None of them is absolute, but a project that meets most of them is a reasonable candidate. · The project needs a small, well-defined amount of money for direct costs, typically in the range of a few hundred to a few tens of thousands of dollars. · The question is genuinely exploratory: it is worth asking even though the answer is uncertain, and conventional funders are unlikely to support it until someone has taken a first look. · The research can be explained to a curious non-specialist in a few sentences, and the connection between the money and the answer is concrete. · The researcher, or someone on the team, is willing to spend real time on communication: before the campaign, during it and for the whole life of the project. · The research can be carried out within any required ethical and regulatory oversight, and the researcher knows what that oversight is. · The researcher is prepared to share the results publicly, including results that are negative, disappointing or ambiguous. A project that fails several of these tests may still be worth doing, but it is probably not worth crowdfunding. The next chapter is about how to take an idea that almost meets them and shape it so that it does. Chapter 2: Shaping a Project the Crowd Can Back Most crowdfunding campaigns are decided before they launch, and most of that decision happens when the researcher chooses what, exactly, to ask the crowd to fund. A good research idea is not automatically a good crowdfunding project. The same underlying question can be framed as a campaign that is vague, expensive and open-ended, or as one that is specific, affordable and finished within months. The second will raise money more easily, and it will also be far easier to deliver honestly. This chapter is about the craft of turning a research idea into that second kind of project. The unit of funding Grant proposals typically describe a programme of work: several aims, spread over several years, connected by a larger hypothesis. That structure makes sense when a panel is deciding whether to invest in a research direction. It makes little sense for a crowd. A backer is not investing in a research direction. She is paying for a specific piece of work to happen, and she wants to know what that piece of work is, what it will cost, and what she will learn at the end. The right unit of funding for a crowdfunding campaign is therefore not a programme but a step: one discrete, well-bounded activity that produces one identifiable output. Consider the difference between two framings of the same idea. A marine biologist interested in how microplastics affect filter-feeding invertebrates might describe her research as "understanding the impacts of microplastic pollution on coastal ecosystems." That is a fine description of a career. As a crowdfunding pitch, it is hopeless: it has no end, no clear cost, and no obvious output. Alternatively, she might propose to "collect mussels from five sites along a single estuary, from the harbour mouth to the upper reaches, and count the microplastic particles in their tissue, to find out whether contamination falls off with distance from the city." That is a project. It has a place, a method, a sample size, a result that will exist in a few months, and an answer that anyone can understand, whichever way it comes out. The shift from programme to step is the single most important move in designing a crowdfunded project. It has several consequences. It forces the researcher to identify what is genuinely the first thing to do. Many large research ideas contain a small, cheap experiment that would tell the researcher whether the rest is worth pursuing. Crowdfunding is ideal for exactly that experiment. It makes the budget concrete. A step can be costed item by item. A programme can only be estimated. It makes success and failure legible. At the end of a step, the researcher can say what was done and what was found. At the end of a programme, there is always more to do. And it makes the campaign repeatable. A researcher who funds step one through a campaign can report the results to the same backers and, if the results warrant it, launch step two. Several researchers have built sequences of campaigns this way, each funded partly by the backers of the last. The high-risk question Crowdfunding is often described as a way to fund high-risk research, and that is true, but the phrase needs care. There are two very different kinds of risk in research, and they call for different treatment. The first is scientific risk: the possibility that the hypothesis is wrong, that the effect does not exist, that the measurement shows nothing. This kind of risk is the whole point of exploratory research, and it is exactly what conventional funders are reluctant to take on. A crowdfunded project can embrace it openly. Backers can be told, in plain terms, that the answer might be no, and that a no is still worth knowing. Many are delighted to fund a genuine gamble, provided it is described as one. The second is execution risk: the possibility that the researcher cannot carry out the work at all, because the method does not function, the equipment cannot be obtained, the permits are not granted, the samples cannot be collected or the team does not have the skills. This kind of risk is not a feature. A backer who funds an experiment expects the experiment to be done, whatever its result. A project that might never produce any result at all, not because the hypothesis fails but because the work cannot be carried out, is a much weaker candidate for public money. The design goal, then, is to maximize the proportion of risk that is scientific and minimize the proportion that is execution. A project in which the question is uncertain but the method is proven is a good crowdfunding project. A project in which the question is uncertain and the method is also untested is a poor one, unless the untested method is itself the question and success or failure of the method is the result being reported. The Glowing Plant campaign of 2013 is a vivid example of confusing the two. The project's backers were promised seeds of a plant that would glow visibly in the dark. Whether such a plant could be engineered with the tools then available was a genuine scientific question, and a fascinating one. But the campaign framed the outcome as a deliverable, a product that backers would receive, rather than as an experiment that might fail. When the engineering proved much harder than expected, with the introduced genes producing only dim light and the pathway difficult to transfer, the project had no honest way to fulfil its promises. The scientific risk had been sold as an execution certainty. Tests for a crowdfundable project Before committing to a campaign, it helps to put the project through a set of questions. Each is simple; together they catch most of the common failures. Can the question be stated in one sentence that a curious fourteen-year-old would understand? Not simplified to the point of inaccuracy, but stripped of jargon. "Do mussels closer to the city carry more microplastic?" passes. "Characterizing spatial heterogeneity in anthropogenic particulate burden in Mytilus" does not, though it describes the same study. Can the answer be delivered, in some form, within a period backers will tolerate? For most campaigns, that means results within six to eighteen months of funding. Longer projects can be crowdfunded, but they need a plan for intermediate outputs that backers will see along the way. Is there a clear link between the money and the work? A backer should be able to see that her money buys specific things: sequencing, travel, reagents, equipment, participant payments. A budget that consists mainly of salary for an unspecified amount of time is harder to support and harder to account for. Is the answer interesting whichever way it comes out? A project whose only publishable, communicable outcome is a positive result creates pressure to find one. A project whose answer is interesting either way, such as whether contamination declines with distance or does not, removes that pressure and makes honest reporting easier. Can the researcher actually do it? The researcher, or the team, should have the skills, the access to equipment, the permissions and the time. If any of these is uncertain, the uncertainty should either be resolved before launch or stated plainly on the campaign page. Is the project within ethical and regulatory bounds, and are the approvals in hand or clearly obtainable? Research involving human participants, animals, protected species, controlled substances, genetically modified organisms or sensitive data is subject to oversight that does not disappear because the funding source is unusual. Chapter 8 discusses this in detail. The practical point here is that approvals should be planned before launch, not after. Would the researcher be comfortable having every backer read the final report? This is the simplest test of all, and it catches a surprising number of problems. If the honest answer is no, because the project is really a pretext for something else, or because the promised outcomes are exaggerated, or because the researcher does not intend to report negative results, the project should be rethought. Types of project that work well Experience across the research platforms suggests several recurring patterns of project that suit the model. The pilot. A small study designed to generate the preliminary data that a larger grant application needs. Pilots are ideal because they are exactly what conventional funders are least willing to pay for and most want to see. A researcher can say to backers, truthfully, that their money is buying the evidence that may unlock a much larger investment later. The field season. A defined period of data collection in a particular place: surveying a population, sampling a site, recording an event. Field seasons have natural boundaries, concrete costs such as travel, permits and equipment hire, and strong narrative appeal. Backers can follow the expedition in real time. The sample run. A batch of laboratory analysis on material the researcher already has: sequencing, isotope analysis, dating, chemical assays. These projects are cheap to describe, cheap to cost and relatively low in execution risk, because the samples already exist. The instrument or tool. Building a piece of equipment, writing software or developing a method that the researcher, and often others, will use. Tools work well when they will be released openly, since backers can see a lasting public benefit beyond the single study. The community study. Research carried out with and for a particular community, such as patients with a rare disease, residents of a polluted neighbourhood or practitioners of a craft. These projects often have a ready-made audience that cares intensely about the question. They also carry the heaviest ethical responsibilities, discussed in Chapter 8. The replication or re-analysis. Repeating an influential study, or re-analysing public data with a new question. These are chronically underfunded because they carry little prestige, and yet many members of the public find them compelling once the stakes are explained. Types of project that work badly Some projects are poor fits regardless of how well the campaign is run. Anything promising a treatment or cure. Patients and families facing serious illness are among the most generous backers of research, and also among the most vulnerable to overpromising. A campaign that implies, even indirectly, that its results will lead to a treatment for a particular person or disease creates expectations research cannot meet. Campaigns connected to disease can be entirely legitimate, but they must be scrupulous about describing what the work can and cannot achieve. Clinical trials of interventions. Testing a treatment in people requires regulatory approval, trial registration, insurance, safety monitoring and a scale of funding that crowdfunding rarely reaches. There are rare exceptions run by established organizations with full oversight, but for an individual researcher this is not the right tool. Large capital projects. Equipment costing hundreds of thousands of dollars, facilities, long-term salaries. Crowdfunding rarely reaches these sums, and when it does, the obligations to backers become correspondingly heavy. Work that cannot be explained. Some important research is simply too technical to communicate without substantial background. That is no criticism of the research. It means the research should be funded some other way. Projects whose real purpose is a product. A study designed mainly to validate a commercial product, or a campaign whose rewards are the real attraction, is a pre-sale, not research funding. General-purpose platforms serve pre-sales well, but they should be described as such, and the research framing should not be used to lend them credibility. Reshaping an idea Many ideas fail these tests in their first form and pass them after reshaping. The reshaping usually involves some combination of four moves. The first is narrowing. A study of five species becomes a study of one. A survey of a whole region becomes a survey of a single valley. A comparison of ten conditions becomes a comparison of two. The narrowed project is less ambitious but far more deliverable, and the broader version can follow if the first result is interesting. The second is sequencing. The project is divided into stages, and the campaign funds only the first. The researcher explains the full plan so that backers see where the work is heading, but promises only what the first stage will produce. The third is de-risking. The researcher identifies the parts of the plan most likely to fail for practical reasons, and either resolves them before launch, for example by running a quick test of the method or securing the permit, or restructures the project so that their failure still produces a reportable result. The fourth is reframing the output. Instead of promising a discovery, the researcher promises a measurement. Instead of promising a working device, she promises a tested prototype and a report of how it performed. The output is something she controls, the doing of the work and the reporting of what happened, rather than something she does not, the answer nature gives. Consider how these moves might apply to a hypothetical proposal from an independent researcher interested in whether a traditional fermented food has antimicrobial properties. In its first form, the idea is to "discover new antibiotics in traditional foods." That promise is enormous, open-ended and misleading, since the chance of finding a clinically useful antibiotic in any single study is very small. Narrowed, it becomes a study of one food from one region. Sequenced, the first stage is simply culturing the microbes present and testing extracts against a panel of harmless laboratory bacteria. De-risked, the researcher first confirms that she has access to a laboratory with appropriate biosafety arrangements and has run a small trial of the culturing method. Reframed, the promised output is not "new antibiotics" but a public report and open dataset describing which microbes live in the food and whether any of them inhibit the growth of test bacteria in the dish. The reshaped project is smaller, more honest and much more likely to be funded, and its result, whether positive or negative, is something the researcher can deliver. The shape of the promise Every choice in this chapter is, at bottom, a choice about what the researcher will promise. The question, the scope, the output and the timeline together define the promise that backers will hold the researcher to. The best time to get that promise right is before a single dollar has been raised, when changing it costs nothing. A good promise is specific enough that backers know what they are paying for, modest enough that the researcher can keep it whatever nature decides, and interesting enough that people want it kept. The rest of the campaign, from the platform to the budget to the video, is built on that foundation. Chapter 3: Platforms, Models and the Rules of the Game Once a project has been shaped, the researcher faces a set of practical choices that are easy to make carelessly and hard to reverse: which platform to use, which funding model to adopt, whether to run the campaign in her own name or through an institution, and what legal and financial obligations follow. These choices determine who sees the campaign, how the money flows, what happens if the target is not met, and what the researcher owes, formally and informally, to the people who give. This chapter works through them in turn. Four models of crowdfunding The word "crowdfunding" covers several quite different arrangements, distinguished by what backers receive in return for their money. Four are relevant to research. In donation-based crowdfunding, backers give money and receive nothing tangible in return beyond thanks and, usually, updates on how the money is used. This is closest to charitable giving, and when the recipient is a registered charity or a qualifying institution, contributions may be tax-deductible for the donor. Most research-specific platforms operate largely on this model. In reward-based crowdfunding, backers receive something in exchange for their contribution: a postcard from the field site, a named acknowledgment in the eventual paper, a print of a microscope image, a video call with the researcher, a copy of a book. This is the model of the general creative platforms such as Kickstarter. Rewards can boost participation, but they add cost, work and obligations, and they change the legal character of the transaction, since a reward is something closer to a purchase than a gift. In equity crowdfunding, backers receive a share in a company. This is regulated as a securities offering. In the United States, the Jumpstart Our Business Startups Act of 2012 created a framework for it, and the Securities and Exchange Commission's Regulation Crowdfunding rules came into effect in 2016. Equity crowdfunding is relevant only to research carried out inside a company with commercial prospects, and it brings the full apparatus of investor disclosure and financial regulation. In lending or debt-based crowdfunding, backers lend money that is to be repaid, sometimes with interest. This is almost never appropriate for research, since research does not generate the cash flows needed to repay a loan. For the kind of exploratory research this book is about, the relevant models are donation-based and reward-based, and the choice between them has consequences explored below. Table 1 sets out the main differences across all four. Table 1. Crowdfunding models compared for research use. Model What backers receive Typical platforms Main obligations on the researcher Fit for exploratory research Donation-based Thanks, updates and access to results Research platforms such as Experiment; university giving sites; general donation sites Use funds as described; report honestly; observe charity and tax rules where applicable Strong: matches the gift-like nature of basic research Reward-based A tangible or experiential reward scaled to the contribution General creative platforms such as Kickstarter and Indiegogo Deliver rewards as promised; consumer-protection duties; tax on income in many jurisdictions Workable for projects with a natural product or artefact; risky when the research outcome is the reward Equity A share in a company Regulated equity portals Securities disclosure, investor reporting, financial regulation Only for commercial ventures; poor fit for open research Lending Repayment, sometimes with interest Peer-to-peer lending sites Repayment on schedule Almost never appropriate All-or-nothing versus keep-what-you-raise A second dimension cuts across the models: what happens if the campaign does not reach its target. Under an all-or-nothing rule, backers are charged only if the target is met. If the campaign falls short, no money changes hands and the researcher receives nothing. Kickstarter and Experiment both work this way. Under a flexible or keep-what-you-raise rule, the researcher receives whatever has been pledged, whether or not the target is reached. Indiegogo has offered this option, and most general donation platforms work this way by default. The all-or-nothing rule can seem harsh, but it serves both parties well in research. For backers, it guarantees that their money will only be spent if there is enough to do the work. Nobody wants to fund forty percent of a sequencing run. For researchers, it provides a clear deadline and a clear stake, and it guards against the uncomfortable position of having raised enough money to be obliged to do something but not enough to do it properly. The rule also has a strategic effect that researchers should understand. Studies of crowdfunding in general, including work by Venkat Kuppuswamy and Barry Bayus on Kickstarter, have found that support tends to cluster at the beginning and end of campaigns, and that backers respond to how close a campaign is to its goal. Under all-or-nothing, the deadline creates urgency, and the approach to the target often draws in final contributions from people who want to push it over the line. This makes the choice of target critically important, a subject taken up in Chapter 4. A flexible rule makes sense only when the project can be scaled smoothly to whatever sum is raised, for example a survey whose sample size can grow or shrink with the budget. Even then, the researcher should state in advance what will be done at different funding levels, so that backers know what their money will buy if the target is missed. Research platforms versus general platforms The choice of platform involves a trade-off between audience and fit. General-purpose platforms such as Kickstarter and Indiegogo attract enormous numbers of visitors and have well-developed tools for campaign pages, rewards and payments. But their audiences are browsing for games, gadgets, films and design objects, and a research project competes for attention with products that are easier to understand and more immediately rewarding. Their rules are also designed for creative projects and products. Kickstarter, for instance, has long required that projects create something to share with others and has restricted certain categories; its prohibition of genetically modified organisms as rewards after the Glowing Plant campaign is one example of how policies can change in response to research projects that sit uneasily with the platform's model. Research-specific platforms offer a better fit. Their pages are designed around the elements a research project needs: a question, methods, a budget, a team, a section for progress notes. Their audiences are self-selected for interest in science. Experiment, the largest, has staff review submitted projects for clarity, scientific accuracy and feasibility before they go live, which provides a modest form of quality control that backers can rely on. Its current terms, as published on its site, are an eight percent platform fee plus payment-processing fees of roughly three to five percent, charged only when a project is fully funded. Some projects on Experiment, those affiliated with qualifying non-profit organizations, are labelled as tax-deductible for United States donors. The trade-off is that research platforms bring relatively little traffic of their own. The evidence reviewed in Chapter 1 suggests that most backers of research campaigns come from the researcher's own networks rather than from browsing the platform. So the choice of platform matters less for finding backers than for how the campaign is presented, how the money is handled and what rules apply. The researcher should expect to bring her own audience wherever she goes. University platforms and institutional routes Many universities now operate their own crowdfunding sites, usually run by the development or advancement office that handles alumni giving. These have particular advantages and disadvantages. On the positive side, money raised through a university site is usually treated as a gift to the institution, which means that donors can typically claim tax deductions where the law allows, the money is held and accounted for by the institution, and the researcher is protected from having the funds treated as personal income. University sites often come with support from communications staff, access to alumni mailing lists and the credibility of the institution's name. On the negative side, university sites are often slower and more bureaucratic. Some restrict who may run a campaign, for example to faculty or to projects approved by a department. Some charge an administrative fee or an overhead levy. They may have limited reach beyond the institution's own supporters, and they may impose restrictions on how projects are described. For researchers at universities, a further question arises even when they use an external platform: whose money is it? Many institutions have policies requiring that funds raised for research conducted by their employees, or using their facilities, be routed through the institution. Some treat crowdfunded money as a gift, with minimal overhead; others treat it as a form of sponsored research, with the full apparatus of indirect cost recovery. A researcher who raises money personally on an external platform for work done in a university laboratory may find herself in breach of institutional policy, or may find that the university claims a share of the funds. The time to discover this is before launch. A short conversation with the institution's research office or development office usually settles the question. Graduate students face a particular version of this issue. Their research is often supervised and conducted in facilities they do not control, and their right to raise money in their own name varies between institutions. Their supervisors should be involved in the decision, both because the supervisor may have obligations to the institution and because the supervisor's support, and network, is often valuable to the campaign. Independent researchers Researchers without institutional affiliation face a different set of issues. They are free of institutional policies, but they lose the protections and infrastructure that institutions provide. Money raised by an individual on a crowdfunding platform is, in many jurisdictions, treated as taxable income to that individual, particularly when rewards are offered, and particularly when the sums are significant. The rules differ widely between countries and are not always clear. In the United States, payment processors report certain transactions to the tax authority, and the tax treatment of crowdfunding receipts depends on whether they are gifts, income or something else. An independent researcher should seek advice from an accountant familiar with the relevant jurisdiction before launching a campaign of any size, and should build the expected tax into the budget. Some independent researchers address these issues by working through a fiscal sponsor: a registered non-profit that accepts funds on behalf of a project and administers them for a fee. Fiscal sponsorship can make donations tax-deductible, provide financial accountability and give the project a legal home. Some independent research organizations have been established partly to serve this function for researchers outside universities. Independent researchers also need to arrange for the facilities, oversight and insurance that institutions ordinarily supply. A study involving human participants needs ethical review regardless of who funds it, and an independent researcher may have to use a commercial or independent review board. Laboratory work may require access to rented facilities, community laboratories or collaborations with institutions. These arrangements have costs and lead times that belong in the campaign plan. Legal and consumer-protection obligations Crowdfunding sits in a patchwork of law that varies between countries and is still developing. A few general points apply widely. Platforms set terms of use that bind campaign creators. Kickstarter's terms, for example, require creators who cannot fulfil their promises to make a good-faith effort to complete the project, communicate honestly with backers about what went wrong, account for how the money was spent and offer refunds where appropriate. These terms are contractual. They are enforceable in principle, even if enforcement is rare. Consumer-protection law may also apply, particularly to reward-based campaigns. In 2015 the United States Federal Trade Commission brought its first action against a crowdfunding creator, a man who had raised money on Kickstarter for a board game, failed to deliver it and spent much of the money on personal expenses. The case established that regulators regard crowdfunding promises as potentially enforceable representations to consumers. State attorneys general have brought similar actions. Research campaigns rarely attract this kind of scrutiny, but the principle is clear: promises made on a campaign page can have legal weight. Charity law applies when donations flow through registered charities, including universities and fiscal sponsors. Funds given for a specified purpose may be legally restricted to that purpose. If the project cannot be carried out, the charity may be obliged to seek donors' consent before redirecting the money, or to return it. None of this should deter a researcher from crowdfunding. The obligations it imposes are, for the most part, obligations an honest researcher would honour anyway: do what you said, tell people what happened, account for the money and offer a remedy if things go wrong. But it does mean that the promises on a campaign page should be written with care, and that a researcher should know which rules govern her particular campaign before it launches. Making the choice Pulling this together, most researchers will find that their situation points fairly clearly to one route. A university researcher with a modest project and a supportive institution will often do best on the institution's own platform, or on a research platform with the funds routed through the institution, after checking policy with the research office. A graduate student with a well-defined project and a supervisor's support will often do well on a research platform, with the supervisor's agreement and an understanding of how the funds will be held. An independent researcher will usually need either a fiscal sponsor or a clear understanding of the personal tax consequences, and will need to budget for the review, facilities and insurance that an institution would otherwise supply. A researcher whose project has a natural, deliverable artefact, such as a book, a tool, a dataset with public value or a documentary, may reasonably choose a general reward-based platform, provided she keeps a sharp line between the reward and the research outcome. In every case, the choice should be made with the promise in mind. The platform, the model and the institutional route are the vehicles for that promise, and each shapes how it will be understood and enforced. Hashtags: #ScientificCrowdfunding #ResearchCrowdfunding #PublicScience #CommunityBackedResearch #ResearchCampaigns #CrowdfundingPlatforms #ExperimentPlatform #DonationBasedCrowdfunding #RewardBasedCrowdfunding #AllOrNothingFunding #ResearchCommunication #ScienceOutreach #PublicEngagement #BackerCommunity #CampaignDesign #ResearchBudgeting #CampaignPitch #PitchVideo #AudienceBuilding #PilotResearch #ExploratoryResearch #OpenResearch #ResearchTransparency #CrowdfundedScienceEthics #FutureOfScientificCrowdfunding

  • Research Commercialization (From Academic Patent to Deep-Tech Spinout)

    Download the Book (PDF): Introduction Somewhere in almost every research university there is a freezer, a server or a filing cabinet holding an invention that works. It has been published. It has been cited. It may have been presented to a room of industry visitors who nodded, took a business card and never called. The inventor knows it could matter outside the lab: a diagnostic that would catch a disease earlier, a catalyst that would make a chemical process cheaper, a material that would let a battery survive more cycles. And yet nothing happens. The paper goes into the literature, the graduate student who did the work graduates, and the invention settles into the long sleep of things that were proved possible and never made real. The gap between a working result in a laboratory and a product that someone will pay for has a name in policy circles: the valley of death. The phrase is dramatic, and deliberately so. It describes a stretch of development where the work is too applied for most research funders, too risky for most corporate buyers and too early for most investors. Public money pays for discovery. Private money pays for scale. In between lies a zone where nobody's job description obliges them to pay, and where most promising technologies quietly die. This book is for the academic who wants to get an invention across that zone. It is written for the principal investigator with a result she believes in, the postdoctoral researcher weighing whether to found a company rather than chase a faculty post, and the graduate student who has noticed that the thing they built for their thesis is better than anything on the market. It is also written for the people around them: department heads, technology transfer staff, and the early investors and advisers who have to work with scientists and are sometimes baffled by how differently scientists think. The argument of this book The central claim of what follows is simple, and it explains most of the frustrations academic inventors experience. The valley of death is not primarily a shortage of money. It is a shortage of the right kind of evidence. A scientist spends a career learning to produce one kind of proof: evidence that a phenomenon is real, that a mechanism works as proposed, that a result is reproducible and novel. That is the evidence peer reviewers demand and grant panels reward. But every audience on the far side of the valley wants a different kind of proof. A patent examiner wants evidence that the invention is new, not obvious, and described well enough for someone else to make it. A licensing officer wants evidence that a company somewhere will pay for the rights. A grant agency funding translation wants evidence that the next experiment will retire a specific technical risk. A venture capitalist wants evidence that the technology can become a business returning many times the money put in, within the life of a fund. A university conflict-of-interest committee wants evidence that the inventor's financial stake will not bend the science, harm students or endanger research subjects. These are not the same evidence, and the gap between them is where inventions stall. A paper showing that a molecule binds its target with remarkable affinity is a triumph in a journal and nearly useless in a pitch meeting, where the questions are about manufacturability, toxicity, the size of the patient population, the regulatory path and whether anyone else owns the chemistry. The inventor who treats each new audience as a harder version of peer review, and responds by generating more of the same evidence, will keep losing. The inventor who learns to see each stage as a translation problem, converting what is known into what the next gatekeeper needs to believe, has a real chance. Seen this way, each of the institutions in this book stops being an obstacle and becomes a converter. The technology transfer office converts a discovery into owned, licensable property. The patent converts an idea into an exclusive right that an investor can value. Translational grants convert technical uncertainty into data. The founding team converts a licence into an organisation that can act. Venture capital converts demonstrated potential into the money needed to realise it. And a well-run conflict-of-interest process converts a dangerous tangle of loyalties into an arrangement that lets the inventor stay a scientist while becoming a founder. None of these conversions happen automatically, and each one has its own grammar. What the book covers, and in what order The chapters follow the rough order in which an academic inventor meets each problem, though in practice they overlap and loop back. The first chapter looks at the valley of death itself: why it exists, what economists and policymakers have said about it, and why deep-technology inventions, those grounded in new science or engineering rather than in new business models, fall into it more often than software or consumer products do. It introduces the idea, which runs through the rest of the book, that each audience on the far side of the valley needs its own kind of evidence. The second chapter turns to the university's technology transfer office, the institution most academic inventors meet first and understand least. It explains who owns an academic invention and why, how the Bayh-Dole Act reshaped university patenting in the United States and influenced policy elsewhere, what an invention disclosure is for, and how to work with a transfer office rather than against it. The third chapter is about patents: what they protect, what they cannot protect, and the specific ways in which academic habits, especially the rush to publish and the conference talk given before a filing, can destroy rights that would otherwise have existed. The fourth chapter moves from a single patent to a portfolio and a licence, explaining how a university decides between licensing to an established company and backing a spinout, and what the terms of a spinout licence typically contain. The fifth chapter deals with non-dilutive funding: the grants, prizes and public programmes that pay for translation without taking equity. These are often the most valuable money a deep-tech company ever receives, and they are badly misunderstood. The sixth chapter covers the founding of the company itself: who should run it, how equity should be divided, what role the academic founder should take and the mistakes that most often poison a young company before it raises its first round. The seventh chapter is about venture capital: how venture funds actually work, why their economics shape every question an investor asks, how to pitch a scientific company to people who are not scientists, and what a term sheet means. The eighth chapter addresses conflicts of interest and commitment, the area where academic inventors most often get into real trouble, and where the stakes include research integrity, student welfare and, in clinical research, human lives. A note on scope A short book on a subject this large has to leave things out. The emphasis here is on the United States, the United Kingdom and continental Europe, because that is where most of the relevant policy and most of the academic spinout activity has been concentrated, though many of the same patterns hold in Canada, Australia, Israel, Singapore and elsewhere. The book concentrates on inventions that need new science or engineering to reach the market, in fields such as therapeutics, diagnostics, medical devices, advanced materials, energy, semiconductors, quantum technologies and robotics. Pure software start-ups founded by academics face some of the same questions, but they rarely depend on patents or on long translational development, and much of this book would matter less to them. The book also assumes a reader who is willing to learn a new language without surrendering the old one. A recurring theme is that the best academic founders do not stop being scientists. They add a second competence alongside the first. They learn to hear what a licensing officer, an examiner or an investor is actually asking, and they answer in terms that audience can use, while protecting the integrity of the research that made the invention possible in the first place. Much of what follows touches on law and finance: ownership of intellectual property, patent procedure, licence terms, share structures, investment agreements and regulations governing financial interests in research. It is offered as general information about how these systems usually work, to help an inventor understand the landscape and ask better questions. It is not legal, patent or investment advice. Rules differ between countries, between institutions and over time, and the specific facts of an invention matter enormously. Before signing, filing or investing anything, an inventor should consult a qualified patent attorney, a lawyer who acts for them personally rather than for the university or an investor, and, where money is involved, a qualified financial adviser. Why this matters It is easy to treat academic commercialisation as a side issue, a way for universities to earn a little licensing income or for a few scientists to become rich. That view misses what is at stake. Many of the technologies that define modern life began in academic laboratories and crossed into industry through the mechanisms this book describes: recombinant DNA, the search algorithm that built Google, the lithium-ion chemistries that power phones and cars, messenger RNA vaccines, CRISPR gene editing and a long list of medicines. Each of those crossings was contingent. Each could have failed at several points, and the history of science is full of inventions of similar promise that did. Every invention that dies in the valley represents public research money that produced knowledge but no benefit anyone could use. Every one that crosses, by contrast, can create jobs, treatments, cleaner industries and the tax revenue that funds the next generation of research. The inventor who learns to cross is doing more than building a company. They are completing the social bargain that justified the research funding in the first place. The path is hard and the odds are poor. Most spinouts fail, as most start-ups do. But failure in the valley is not random. It follows recognisable patterns: rights lost before anyone thought to file, licences negotiated badly, funding sought from the wrong source at the wrong time, founding teams that could not function, pitches that answered the wrong questions, conflicts left unmanaged until they became scandals. Each of these is avoidable, and the chapters that follow explain how. Chapter 1: Why Good Science Dies The phrase "valley of death" entered American science policy in the 1990s and was popularised in Congressional testimony and reports about the gap between federally funded research and commercial development. Its most careful early treatment came in a 2002 study for the National Institute of Standards and Technology by Lewis Branscomb and Philip Auerswald, Between Invention and Innovation. They argued that the image of a valley, a single chasm between two well-funded plateaus, was slightly misleading. The terrain looked more like a sea of uncertainty in which some projects swam and most sank, where survival depended on a patchwork of angel investors, corporate partnerships, government programmes and luck. Yet the core of the metaphor has survived because inventors recognise it. There is a stage of development where nobody is obliged to pay, and it is exactly the stage where an academic invention usually sits. To cross it, it helps to understand why it exists. The valley is not an accident of bad policy that a better funding scheme could simply abolish. It follows from the different logics of the institutions on either side. Two plateaus with different rules On one side sits the research system. Public funders, charities and universities pay for work whose value lies mainly in knowledge. The economist Kenneth Arrow set out the classic reasoning in 1962: knowledge is hard to own, easy to copy once revealed, and uncertain in its payoff, so private markets will produce less of it than society would want. Governments therefore fund basic research directly. The research system rewards novelty, rigour and publication. It is organised around the individual investigator and the grant cycle, which typically runs three to five years. Its outputs are papers, trained people and, occasionally, inventions. On the other side sits the commercial system. Companies and investors pay for work whose value lies in products that customers will buy. They reward predictability, cost control, scale and return on capital. Established companies fund development when the technical risk is modest and the market is clear, because their shareholders expect steady returns and their internal budgets favour projects likely to reach revenue within a few years. Venture investors accept more risk, but they need a credible path to very large returns within the life of a fund, commonly ten years. The valley opens because an academic invention usually fails the entry test of both systems at once. It has already produced its main contribution to knowledge, so research funders are no longer interested in paying for the unglamorous work of making it reliable, manufacturable and cheap. Nobody writes a high-impact paper showing that a promising catalyst still works after ten thousand hours, or that a device can be assembled with a tolerance a factory can hold. Yet that work is exactly what the commercial side needs to see before it will commit money. The invention is too applied for one side and too risky for the other. Technology readiness and the missing middle Engineers describe this gap with technology readiness levels, a scale developed at NASA in the 1970s and 1980s and later adopted by the US Department of Defense, the European Commission's research programmes and many national funders. The scale runs from level 1, where basic principles have been observed, to level 9, where the actual system has been proven in operational use. Most academic research sits between levels 1 and 3: principles observed, a concept formulated, a proof of concept shown experimentally. Most commercial investment begins around levels 6 or 7, where a prototype has been demonstrated in a relevant or operational environment. Levels 4 and 5, validation in the laboratory and then in a relevant environment, are the floor of the valley. This is where the question changes from "does it work?" to "does it work reliably, outside the conditions that made it work the first time, at a cost and scale that matter?" It is expensive, slow, intellectually unfashionable and full of unpleasant surprises. It is also where the value of the invention is actually determined. The readiness scale was designed for engineered systems, and it maps imperfectly onto drugs, where the milestones are set by regulators and clinical trials, or onto materials, where scale-up often reveals entirely new chemistry. But the underlying point generalises. There is a middle phase of development in which technical risk is still high and the evidence that would persuade a commercial buyer does not yet exist, and nobody on either plateau sees paying for it as their job. Why deep technology falls further The valley is not equally deep for every kind of invention. A new mobile application can go from idea to paying customers in months with a few hundred thousand dollars. Its technical risk is low, because the tools are well understood, and its main uncertainty is whether customers want it. A new battery chemistry, cancer therapy, semiconductor process or fusion component faces technical risk, market risk and often regulatory risk all at once, and it needs years and tens or hundreds of millions of dollars to resolve them. The term deep tech, which became common in the 2010s, describes this second class: ventures built on substantial scientific or engineering advances rather than on new business models applied to existing technology. Deep-tech ventures share several features that make the valley wider. They need capital before they have customers. A software company can often sell a rough version of its product early and use revenue to fund development. A company developing a new medicine may spend a decade and hundreds of millions of dollars before selling anything at all. They carry technical risk that cannot be resolved by talking to customers. Customer discovery tells you whether people want the thing; it cannot tell you whether the physics will cooperate at scale. Some questions can only be answered by building, and building is expensive. They often require specialised infrastructure: clean rooms, pilot plants, animal facilities, clinical sites, regulatory expertise. These are costly to build and hard to rent. Their markets are frequently dominated by large incumbents who control distribution, manufacturing and customer relationships. A new diagnostic may be scientifically superior and still fail because hospital purchasing, reimbursement codes and clinical guidelines are all built around the existing test. And their timelines are long. Venture funds are usually structured to return money within about ten years, with most investments made in the first few. A technology that needs twelve years to reach meaningful revenue is awkward for a fund of that shape, regardless of its eventual value. None of this makes deep tech a bad investment. Some of the most valuable companies of the last half century, from Genentech to the major semiconductor equipment makers, are deep-tech companies. But it means the evidence a deep-tech inventor needs to produce before anyone will fund them is harder, slower and more expensive to generate than the evidence a software founder needs, and the stretch of development during which nobody wants to pay for it is longer. The evidence problem The claim of this book is that the valley is at heart an evidence problem. Money follows evidence; the reason nobody pays for the middle phase is that nobody yet has the evidence that would justify paying, and generating that evidence is itself the thing that costs money. Breaking the circle requires seeing clearly what evidence each audience needs. An academic career trains one kind of evidence very well. A good scientist knows how to design a controlled experiment, how to rule out alternative explanations, how to show that a result is statistically robust and how to write it up so that peers accept it. These skills remain essential for commercialisation, because an invention that is not real cannot be rescued by any business plan. But they are only the first rung. The table below sets out, in simplified form, what each of the main audiences an academic inventor meets is really trying to establish, and the kind of evidence that persuades them. The categories overlap and the details vary by field, but the pattern matters more than any single cell. Table 1 compares these audiences. Table 1. What different audiences need to believe. Audience Core question Evidence that persuades Evidence that does not Journal reviewer Is it true and new? Controlled experiments, reproducibility, novelty against literature Market size, cost estimates Patent examiner Is it new, non-obvious and fully described? Comparison with prior art, enabling description, working examples Citation counts, impact factor Licensing officer Will anyone pay for rights? Named industry interest, defensible claims, a plausible product Elegance of mechanism Translational funder Will this project retire a defined risk? Milestones, go/no-go criteria, a credible team Broad future applications Venture investor Can this return many times the fund's money? Large market, strong team, protected position, path to scale A single impressive dataset Conflict committee Will financial interest bend the research or harm people? Disclosure, separation of roles, independent oversight Assurances of good intent The last column is where academic inventors lose time. Faced with an unpersuaded investor, many respond by producing more of the evidence they already know how to produce: another figure, a higher-resolution image, a tighter error bar. It rarely works, because the investor was not doubting the science. They were doubting whether anyone would buy the product, whether the team could build a company, or whether a larger competitor already had a patent covering the same ground. How inventions actually die It is worth being specific about the ways academic inventions fail in the valley, because they are more varied than the image of a single funding gap suggests. Several patterns recur. Some die before anyone realises they could be commercialised. The inventor publishes, the rights are lost, and by the time someone sees the potential the knowledge is in the public domain. This is not always bad, since open knowledge has great value, but when a product needed private investment to reach users, the loss of patent protection can mean the product never gets made. Some die in the technology transfer office. The office, faced with more disclosures than it can patent, decides the invention is too early or the market too uncertain, and declines to file. The inventor, not understanding why or what might change the decision, gives up. Some die in a licence. The rights go to an established company that has no strong incentive to develop them, perhaps because the technology competes with its existing products or because internal priorities shift. The invention sits on a shelf, protected by patents that keep anyone else from developing it. Some die in the lab, during the middle phase, because the technology simply does not scale. The effect that was dramatic in a small sample vanishes in a larger one; the material that performed beautifully in grams cracks in kilograms; the drug that worked in mice is toxic in larger animals. This is honest failure, and it is the valley working as it should: most early ideas do not survive contact with reality, and finding that out quickly and cheaply is a success of sorts. Some die in the company, from causes that have nothing to do with the technology: a founding team that cannot agree, an equity split that leaves nobody motivated, a chief executive who cannot raise money, a board that runs out of patience. Investors routinely say they back teams more than technologies, and early failure in deep-tech companies is often a team failure. And some die of conflict. A professor's company and laboratory become entangled; students are drawn into unpaid company work; a clinical trial is compromised by a financial interest; a public scandal forces the university to withdraw. These are the rarest failures and the most damaging, because they harm people beyond the venture itself. Each of these failure modes has its own remedy, and most of the rest of this book is organised around them. But the common thread is that at each point an invention needed evidence it did not have, or needed its evidence translated for an audience that could not read it. Timescales, honestly stated Academic inventors routinely underestimate how long commercialisation takes. The misjudgement comes partly from optimism and partly from the grant cycle, which trains people to think in three-year blocks. In therapeutics, the path from a promising academic discovery to an approved medicine commonly takes well over a decade, with clinical development alone often running seven years or more, and the great majority of candidates that enter human trials never reach approval. In medical devices the timelines are usually shorter, but regulatory clearance, clinical evidence and reimbursement can still take many years. In advanced materials and energy, the gap between a laboratory demonstration and a product at industrial scale has historically been measured in decades; lithium-ion batteries, for instance, took roughly two decades from the foundational chemistry of the 1970s and early 1980s to commercial cells in 1991, and far longer to reach the scale and cost that made electric vehicles practical. Semiconductor process innovations often need a similar arc from laboratory to fabrication plant. This does not mean the academic founder must wait fifteen years for any reward. Companies can be sold, licensed or partnered long before a final product exists, and investors are frequently repaid when a larger company acquires a spinout after it has passed a key milestone. But it does mean that the founding decision should be made with a realistic sense of the commitment involved. Founding a deep-tech company is not a sabbatical project. Where the crossing points are If the valley is an evidence problem, then crossing it is a matter of generating the right evidence in the right order, as cheaply as possible, and finding people willing to pay for each step. Historically, several institutions have grown up to do exactly that, and each of them is the subject of a later chapter. University technology transfer offices exist to capture and license the intellectual property that research produces, and to steer inventions towards companies that can develop them. Patents turn inventions into assets that investors can value and that protect the long, expensive development needed to cross the valley. Translational and small-business grants, from programmes such as the US Small Business Innovation Research scheme, the European Innovation Council, national innovation agencies and medical research charities, pay for the middle phase without taking ownership. Venture capital, especially the subset of firms that specialise in deep technology, takes on high risk in exchange for a share of the upside. And spinout companies are themselves a mechanism for crossing, because they create an organisation whose only job is to push one technology forward, unlike a university laboratory, which must keep producing new knowledge, or a large company, which must balance many priorities. No single one of these institutions spans the valley. The successful crossings are almost always assembled from several of them in sequence: a disclosure leads to a patent, the patent supports a translational grant, the grant produces data, the data support a licence to a new company, the company raises seed money, the seed money funds a demonstration, and the demonstration attracts a larger round. At each handoff there is a translation task. The inventor who understands what each institution needs, and why, can manage those handoffs far better than one who treats each as an unwelcome bureaucratic hurdle. A different way of seeing the inventor's job The most important shift an academic inventor can make is to stop thinking of commercialisation as something that happens after the science is done. In the research system, the paper is the end of a project. In the commercialisation system, the paper is closer to the beginning, and it can even be an obstacle, if it disclosed the invention before rights were protected or oversold results that later work cannot reproduce. The shift is not towards being less rigorous. Investors and partners are, if anything, more unforgiving of irreproducible results than journals, because they lose money on them. A 2011 report from scientists at Bayer and a 2012 commentary by researchers at Amgen both described how often industry teams failed to reproduce published preclinical findings, and those reports sharpened a lesson experienced investors had already learned: that a striking academic result is a hypothesis, not a product. The inventor who welcomes independent replication, runs blinded validation and publishes negative results about their own technology's limits earns a credibility that no pitch deck can buy. The shift is towards a wider view of what counts as evidence, and towards treating every stage as a chance to generate the proof the next audience will need. That begins before the first patent is filed, with the inventor's relationship to the institution that almost certainly owns the invention, whether the inventor realises it or not. Chapter 2: Who Owns the Invention, and the Office That Manages It Many academic inventors are surprised, sometimes angry, to learn that the invention they conceived, built and published does not belong to them. In most research universities in the United States, the United Kingdom and much of Europe, an invention made by an employee in the course of their research, using university resources, belongs to the university. The inventor is named on the patent, shares in any income and often has a strong say in what happens next. But the legal owner, the party with the right to license, sell or abandon the invention, is usually the institution. Understanding why that is, and how the office that manages university inventions actually works, is the first practical step in commercialisation. A remarkable amount of friction between inventors and universities comes from mismatched expectations rather than genuine conflict. How universities came to own inventions In the United States, the modern system was shaped by the Patent and Trademark Law Amendments Act of 1980, universally known as the Bayh-Dole Act after its Senate sponsors, Birch Bayh and Bob Dole. Before 1980, inventions arising from federally funded research generally belonged to the federal government unless an agency granted a waiver, and the government licensed them non-exclusively. Few companies were willing to invest in developing a technology that any competitor could license on the same terms, and a large proportion of government-owned patents went unused. Bayh-Dole allowed universities, small businesses and non-profit institutions to elect to retain title to inventions made with federal funding, on conditions. The institution must disclose the invention to the funding agency, decide within a set period whether to take title, file patent applications in a timely way, try to commercialise the invention, give preference to small businesses when licensing, and share royalties with the inventor. The government keeps a non-exclusive licence to use the invention for its own purposes and holds so-called march-in rights, allowing it to require further licensing if the owner fails to take effective steps towards practical application or in certain other circumstances. March-in has been petitioned many times, especially over drug prices, but no agency has ever exercised it. The law is widely credited with, and sometimes blamed for, the growth of university patenting and licensing since 1980. Before it, only a handful of universities ran serious patent operations; the Wisconsin Alumni Research Foundation, founded in 1925 to manage Harry Steenbock's vitamin D patents, and Research Corporation, founded in 1912, were notable early exceptions. Afterwards, nearly every research university created a technology transfer office. Scholars have argued about how much of the change was caused by Bayh-Dole and how much by concurrent developments, including the biotechnology revolution and a 1980 Supreme Court decision, Diamond v. Chakrabarty, that allowed patents on genetically engineered organisms. But the institutional model it established, in which universities own and license academic inventions and share income with inventors, has been imitated in various forms across much of the world. Bayh-Dole does not itself hand an invention to the university. In 2011, in Stanford v. Roche, the US Supreme Court held that the Act does not automatically vest title in the institution; the inventor owns their invention in the first instance, and the university obtains title through an assignment from the inventor. That is why university employment agreements and intellectual property policies matter so much. After the decision, many universities revised their agreements to use present-tense assignment language, stating that the employee "hereby assigns" future inventions, rather than a mere promise that they "agree to assign", because the case turned partly on that distinction. Ownership outside the United States In the United Kingdom, inventions made by employees in the course of their normal duties belong to the employer under the Patents Act 1977, and universities generally own inventions made by their academic staff. Students are a different matter, since they are not employees, and university regulations often require them to assign rights when they work on funded projects or use university facilities. In Germany, a long tradition known as the professor's privilege gave university professors ownership of their own inventions, but it was abolished in 2002, and German universities now claim inventions under the Employee Inventions Act. Sweden, by contrast, retains a form of the professor's privilege: Swedish university teachers generally own their own inventions, a notable exception among major research economies. Other countries, including Italy, have moved in both directions over the years. The details differ, and every institution's policy says something slightly different about students, visiting researchers, consulting work, software, copyright and materials. The first practical step for any academic inventor is to read their own institution's intellectual property policy, their employment contract, and the terms of any grant or industry agreement that funded the work. Grants and sponsored research agreements frequently contain their own provisions on ownership and licensing, and a company that funded part of the research may already hold an option to license the result. The invention disclosure The invention disclosure is the formal document by which an inventor tells the university about an invention. It is usually a form, submitted to the technology transfer office, asking for a description of the invention, the names of all contributors, the funding sources, any planned or past publications and presentations, and any known commercial interest. Inventors often treat the disclosure as paperwork. It is better understood as the first commercial document in the life of the invention, and a well-written one changes what happens next. The transfer office uses it to decide whether to spend money on a patent application, and in a busy office that decision is made quickly, often by a licensing officer who is not an expert in the inventor's field and who is managing dozens of other cases. A disclosure that helps its own cause has several features. It states plainly what the invention does that existing approaches cannot, in terms a non-specialist can follow. It distinguishes the invention from the closest known work, including the inventor's own earlier publications. It identifies the evidence that the invention works and is honest about what has not yet been shown. It names the likely applications and, where possible, specific companies or kinds of companies that might be interested, including any that have already expressed interest. And it gives accurate dates for any public disclosure, past or planned. The list of inventors deserves care. Inventorship under patent law is not the same as authorship on a paper. A person is an inventor only if they contributed to the conception of at least one claimed element of the invention; a technician who carried out experiments under instruction, or a senior colleague who secured the funding, is not an inventor merely for that reason. Getting inventorship wrong, in either direction, can make a patent vulnerable, and disputes over it are a common source of bitterness in laboratories. The honest course is to describe each person's contribution accurately and let the patent attorney apply the legal test to the claims as finally drafted. Timing matters as well. The disclosure should be filed well before any planned publication, conference talk, poster, thesis deposit or even a detailed seminar to an outside audience. The reasons are explained in the next chapter, but the short version is that in most of the world, a public disclosure before a patent filing destroys the right to patent. Students, collaborators and shared inventions Ownership becomes complicated as soon as more than one institution, or more than one kind of person, is involved, and academic inventions very often involve both. A graduate student funded by a fellowship rather than a salary may not be an employee at all, and whether the university owns their contribution depends on what they signed when they enrolled or joined a funded project. A visiting researcher may remain employed by a home institution that has its own claim. A collaborator at another university almost certainly does. An industry scientist seconded to the laboratory brings their employer's claims with them. When an invention has contributors from several institutions, each institution typically owns an undivided share, and the institutions must agree which of them will lead on patenting and licensing. These inter-institutional agreements are routine, but they take time, and an invention cannot easily be licensed to a spinout until they are in place. Under US law, each co-owner of a patent may in principle exploit it without the consent of the others, which is precisely why a prospective licensee will insist that all owners join in, or that one has authority to act for all. The practical lesson is to establish, at the start of any collaboration, who is contributing what and under which agreements. It is far easier to settle ownership when an invention is hypothetical than when it has become valuable. Students deserve particular care. They are less able to protect their own interests than faculty, their contribution can be substantial, and a dispute over whether a student is an inventor, or whether they can use their own thesis work in a later company, can damage both the student's career and the venture. A laboratory head who makes the rules explicit protects everyone, including themselves. What a transfer office actually does Technology transfer offices go by many names: tech transfer offices, innovation offices, commercialisation units, knowledge exchange teams. Some are departments within the university; others, such as Oxford University Innovation or Imperial College's former commercialisation arm, have been set up as wholly owned companies. Some universities have outsourced the function or shared it among several institutions. Their core work, however, is similar. The office receives disclosures and assesses them. It decides which inventions to protect, commissions patent attorneys to draft and file applications, and manages the resulting portfolio, paying renewal fees and responding to examiners. It markets inventions to potential licensees, negotiates licence and option agreements, and, increasingly, helps to form spinout companies. It manages material transfer agreements, confidentiality agreements and parts of sponsored research contracts. It collects and distributes licensing income. And, in many universities, it runs or supports entrepreneurship programmes, proof-of-concept funds and relationships with investors. The office usually does this with a small staff. A licensing officer may carry a portfolio of many active inventions, and patent budgets are always limited. Most offices therefore apply some form of triage. They file provisional or priority applications on a relatively wide range of inventions, which is cheap, and then decide within a year which of them merit the much larger costs of international filing. Why good inventions get declined A transfer office declining to file, or dropping an application after the first year, is one of the most demoralising experiences in academic commercialisation. It helps to understand the reasons. The most common is that nobody appears likely to pay. A patent costs money to obtain and more to maintain, and a university cannot recover that cost unless someone takes a licence. If the office cannot identify a plausible licensee, whether an existing company or a spinout with a credible team, it may reasonably conclude that the money is better spent elsewhere. Another is that the invention is too early. A promising observation that has not yet been shown to work in a useful setting may be patentable, but the claims will be narrow or speculative, and the patent may expire before a product exists. Offices sometimes advise waiting until the data are stronger, provided the inventor can hold off publishing. A third is that the claims would be weak: the prior art is crowded, the invention is an incremental improvement, or it falls into a category where patents are hard to obtain or easy to design around. A fourth is that the rights are encumbered. An industry sponsor may hold an option, a funder may impose conditions, or the invention may depend on another party's patents. And sometimes the office is simply wrong, or overwhelmed, or lacks expertise in the field. It happens. But an inventor who responds to a decline by complaining rarely changes the outcome. One who responds by addressing the stated reason often does. If the concern is market interest, a letter from a company saying it wants to evaluate the technology changes the calculation. If the concern is prematurity, a specific experiment that would demonstrate utility, and a commitment to delay publication until it is done, may persuade the office to file. If the concern is cost, a translational grant or an early-stage investor willing to cover patent expenses can shift the balance. Where the university decides not to pursue an invention, many policies allow the rights to be released to the inventor, subject to conditions such as repayment of costs and, for federally funded inventions in the United States, approval from the funding agency. Inventors in this situation should ask. Revenue sharing and what it signals Universities share licensing income with inventors, and Bayh-Dole requires this for federally funded inventions without specifying the proportion. The formulas vary widely. Some institutions pay inventors a third of net income; some use sliding scales that give inventors a larger share of the first slice of income and a smaller share thereafter; some split the remainder between the inventor's department, their laboratory and the central university. Net income usually means income after the office recovers its patent costs, which can take some time. The sharing formula matters less than inventors tend to think, for a simple reason: most inventions never generate significant income. Licensing income across the university sector is highly skewed, with a small number of blockbuster inventions accounting for a large share of the total. The Cohen-Boyer recombinant DNA patents, licensed non-exclusively by Stanford and the University of California from 1980 until they expired in the late 1990s, generated about a quarter of a billion dollars. Stanford's licence of the PageRank search algorithm to Google, in exchange partly for shares, yielded shares that the university sold in 2005 for about 336 million dollars. Such cases are exceptional, and many transfer offices do not cover their own operating costs from licensing income. What the revenue formula does signal is how the university thinks about the relationship. An office that treats licensing income as a central source of revenue will negotiate differently from one that treats knowledge transfer as a public mission and accepts modest returns in exchange for more inventions reaching use. Both views exist, often in the same institution, and inventors benefit from knowing which one they are dealing with. Working with the office rather than against it The relationship between inventor and transfer office works best when each side understands the other's constraints. The office has a limited budget, legal obligations to funders and the university, and a portfolio of competing cases. The inventor has deep technical knowledge, networks in the field and usually the strongest motivation to see the invention used. Several practices make the relationship productive. Disclose early, well before publication, so that the office has time to act. Keep the office informed about planned talks, papers and collaborations; an office that is surprised by a conference abstract cannot protect the invention. Bring market intelligence: the names of people in industry who have asked about the work, the companies that dominate the relevant market, the problems customers complain about. Respond quickly to questions from patent attorneys, since delays can cost priority. And be honest about limitations, because an office that discovers late that a key result was not reproducible will be warier the next time. It also helps to understand the language of the office. A licence grants rights to use a patent or other intellectual property, which may be exclusive, meaning only the licensee may use it, or non-exclusive, and may be limited to particular fields of use, territories or periods. An option gives a company the right, for a fee and a limited time, to evaluate a technology and then negotiate a licence on terms defined in advance. A material transfer agreement governs the sharing of physical research materials, such as cell lines or compounds, and often contains clauses about ownership of any resulting inventions. A confidentiality or non-disclosure agreement allows the invention to be discussed with a company without destroying patent rights, though confidential discussions should still be kept to the minimum needed. When the relationship does break down, the most useful escalation routes are usually internal: the director of the office, a faculty committee on intellectual property, or the vice-president or pro-vice-chancellor responsible for research. Universities vary a great deal in how responsive they are, and in some countries governments have pressed them to be faster and more generous with spinouts. But the inventor who arrives with a clear account of what they want, why, and what evidence supports it will usually get further than one who arrives with a grievance. The office's first real decision concerns whether to protect the invention at all, and on what terms. That decision turns on patent law, which is the subject of the next chapter. Chapter 3: Patents for Scientists A patent is a strange instrument, and scientists often misunderstand it in ways that cost them dearly. It is not a certificate of scientific merit, and it does not give its owner the right to make or sell anything. It is a time-limited right to exclude others from making, using, selling or importing the claimed invention in the country that granted it, in exchange for a public description of how the invention works. The bargain is disclosure for exclusivity: the inventor teaches the public, and in return the state lets them stop others from using the teaching for about twenty years. For a deep-tech venture, that right to exclude is often the single most important asset in the company. It is what allows an investor to believe that if the technology works, a larger competitor cannot simply copy it after the start-up has spent years and many millions of dollars proving it. Without that belief, the long and expensive work of crossing the valley is very hard to finance. Understanding patents is therefore not a legal nicety for an academic founder. It is part of understanding the business. What can be patented Most patent systems require that an invention be patentable subject matter, new, inventive and capable of industrial application or, in US terms, useful. It must also be described clearly and completely enough that a skilled person could carry it out. The subject-matter requirement excludes certain categories. Laws of nature, natural phenomena and abstract ideas cannot be patented as such in the United States, and the European Patent Convention excludes discoveries, scientific theories, mathematical methods and computer programs "as such", along with methods of treatment of the human body. These exclusions matter a great deal to academic inventors, because much academic research consists precisely of discovering natural phenomena. A series of US Supreme Court decisions sharpened the line. In Mayo v. Prometheus in 2012, the Court held that a method of adjusting drug dosage based on the level of a metabolite in the blood was an unpatentable application of a natural law. In Association for Molecular Pathology v. Myriad Genetics in 2013, it held that naturally occurring DNA sequences, even when isolated, could not be patented, although synthetic complementary DNA could. In Alice v. CLS Bank in 2014, it tightened the treatment of software and business methods. The practical consequence for academic founders, especially in diagnostics and computational methods, is that a discovery about how nature works must usually be tied to a specific, concrete application that goes beyond the discovery itself. Novelty means the invention has not been disclosed to the public anywhere in the world before the relevant date, in any form: a paper, a poster, a thesis on a library shelf, a conference talk, a preprint, a website, a product on sale, even a sufficiently detailed conversation with someone not bound by confidentiality. Inventive step, or non-obviousness, means that the invention would not have been obvious to a person skilled in the field, given everything that was already known. Much patent prosecution consists of arguing about this second test. The publication problem The single most damaging habit academic inventors bring to patenting is publishing first. It is also the most understandable, because publication is how academic careers are built, and delaying a paper can mean being scooped. The rules differ between jurisdictions in a way that matters enormously. The European Patent Convention, like many national systems, applies what is called absolute novelty: any public disclosure before the filing date counts as prior art, including the inventor's own. There are narrow exceptions, for instance for disclosures made in breach of confidence or at certain officially recognised international exhibitions, but a journal article, conference talk or preprint published before filing will generally destroy the right to a European patent. The United States is more forgiving. Since the America Invents Act moved the country to a first-inventor-to-file system for applications filed from 16 March 2013, US law has provided a one-year grace period for disclosures made by the inventor, or by others who obtained the information from the inventor. An inventor who publishes can therefore still file a US application within twelve months. Japan, South Korea, Canada, Australia and several other countries also have grace periods, though their terms differ. But relying on a grace period is dangerous. It preserves rights only in the countries that offer it, which often excludes Europe and China, and it leaves the inventor exposed if someone else publishes or files on a related idea in the meantime. A grace period is a safety net for mistakes, not a strategy. The sensible practice is to file first and publish second. Once a priority application has been filed, the inventor can usually publish immediately, since the filing date fixes the novelty date for everything described in it. The delay needed is often only a few weeks, provided the inventor warns the transfer office early. The damage comes from last-minute surprises: an abstract submitted for a conference in three weeks' time, a thesis defence open to the public, a preprint posted on impulse. Two further traps deserve mention. First, what counts as a public disclosure is broad. A poster at an internal departmental day with outside visitors, a grant application later published in summary, a slide deck posted online, or a talk to an industry audience without a confidentiality agreement may all qualify. Second, the priority application protects only what it actually describes. If the inventor files on one version of the invention and then publishes additional data or variations not included in the filing, those additions are unprotected. The filing must be complete enough to support the claims that will eventually matter. The anatomy of a patent A patent document has two parts that matter most: the description, sometimes called the specification, and the claims. The description explains the invention, its background and how to carry it out, with examples. It must be sufficient for a skilled person to reproduce the invention without undue experimentation; in US terms, it must satisfy the enablement and written description requirements. This is where academic data become important. Working examples, experimental results and a range of embodiments give the patent attorney material to support broad claims. A thin description, based on a single experiment, supports only narrow claims. The claims define the legal scope of the right. They are numbered sentences, written in a stylised language, each setting out a combination of features. Anything that includes every feature of a claim infringes it; anything that lacks even one feature does not. The first claim is usually the broadest, and later claims add features that narrow the scope, providing fallback positions if the broad claim is found invalid. Scientists often read a patent by looking at the description, which resembles a paper. Investors and lawyers read it by looking at the claims, because the claims are the property. A patent with a brilliant description and narrow claims may be commercially worthless, since a competitor can design around it by changing one feature. A patent whose claims cover the only practical way to achieve a commercially important result can underpin an entire industry. Claims come in different types. Composition or product claims cover a thing, such as a molecule, a material or a device. Method claims cover a process of making or using something. Use claims, available in some jurisdictions, cover a new use of a known substance, such as a new medical indication for an existing drug. Composition claims are usually the most valuable, because they cover the product however it is made or used, but they are also the hardest to obtain when the thing itself is already known. The filing sequence Patent filing is a sequence of decisions spread over several years, each involving more cost. Academic inventors should understand the sequence because the key decisions arrive at predictable times and because the costs escalate sharply. Table 2 sets out the typical route an international patent family takes, using the common pattern of a US provisional or national priority filing followed by an international application under the Patent Cooperation Treaty. Table 2. A typical international filing sequence. Stage Timing from first filing What it does Relative cost Priority filing (e.g. US provisional or UK application) Month 0 Fixes the priority date; not examined in the provisional case Low International (PCT) application By month 12 Preserves the option to file in most countries; search report and opinion Moderate International publication Around month 18 Application becomes public None directly National and regional phase entry Around month 30 (31 at some offices) Separate applications in chosen countries or regions High, rising with translations Examination and grant Often 3 to 6 years from filing Claims negotiated with each patent office High, ongoing Maintenance and renewal Through expiry, usually 20 years from filing Keeps each patent in force Accumulating The first filing is cheap relative to what follows. A US provisional application can be filed for a modest official fee and is not examined; it expires after twelve months, and its only purpose is to establish a priority date for a later full application. Many universities file provisionals liberally, sometimes drafted largely from the inventor's manuscript, as a way to preserve rights before publication. That is better than nothing, but a provisional that simply attaches a paper may not support the claims that later matter, and the priority date only protects what is actually described. The twelve-month deadline is the first major decision point. By then the office must decide whether to file a full application, usually an international application under the Patent Cooperation Treaty, which is administered by the World Intellectual Property Organization and covers over 150 contracting states. The PCT application does not itself produce an international patent; no such thing exists. It preserves the right to seek patents in member countries and provides an international search report and a preliminary opinion on patentability, which is valuable information about the strength of the claims. The second major decision point comes at about thirty months from the priority date, when the application must enter the national or regional phase in each country or region where protection is wanted. This is where costs rise steeply, because each office charges its own fees, many require translations, and each will examine the application separately with local attorneys. A typical decision might be to enter the United States, the European Patent Office, Japan and China, and perhaps a few other markets, depending on where the product will be made and sold. The European Union's Unitary Patent system, which began operating in June 2023 together with the Unified Patent Court, has changed the European step. An applicant who obtains a European patent can now request unitary effect covering the participating EU member states in a single right, rather than validating the patent country by country, and can enforce it in a single court. Not all European countries participate, and the system introduces the risk that a single central challenge could revoke protection across all of them at once, but it has made broad European coverage simpler and, for many applicants, cheaper. Over the life of a patent family, the accumulated cost of filing, prosecution, translation and maintenance across several major jurisdictions commonly reaches the low hundreds of thousands of dollars or more. That is why university offices triage so carefully at the twelve and thirty month points, and why a licensee or spinout is usually expected to take over the costs once a licence is signed. Priority, competition and interference Because most of the world grants patents to the first to file, timing can decide everything when several groups are working on the same problem. The most famous recent example is CRISPR gene editing. The University of California, representing Jennifer Doudna, Emmanuelle Charpentier and colleagues, filed a priority application in May 2012 describing the CRISPR-Cas9 system. The Broad Institute and the Massachusetts Institute of Technology, representing Feng Zhang's group, filed later in 2012 but claimed the use of the system in eukaryotic cells and obtained US patents first through an accelerated examination route. Because the Broad's applications were filed before the America Invents Act changed the rules, the dispute proceeded under the old first-to-invent system, as a series of interference proceedings about who had conceived the invention first. The US Patent Trial and Appeal Board ruled for the Broad on the eukaryotic claims in 2022, and in 2025 the Federal Circuit sent part of that decision back to the board for further consideration of how conception should be assessed. The European picture developed separately, with its own oppositions and revocations. The lesson for academic inventors is not that they should expect a decade of litigation. Most inventions are never contested at all. It is that the filing date and the precise wording of claims can matter as much as the science, and that in a crowded field a few weeks' delay, or a priority document that failed to describe a key embodiment, can determine who owns a platform technology. Freedom to operate A patent gives its owner the right to exclude others. It does not give the owner the right to practise the invention, because doing so may infringe someone else's patent. A new drug formulation may be patented, but if the drug molecule itself is covered by another company's composition patent, the formulation cannot be sold until that patent expires or a licence is obtained. Freedom to operate is the question of whether a product can be made and sold without infringing others' valid patents in the countries that matter. It is a separate question from patentability, and it is answered by a separate analysis: a search of patents in force, followed by legal review of the claims against the planned product. Academic inventors rarely think about it, because research use enjoys some limited protection in some jurisdictions, and because a laboratory is not selling anything. Investors think about it constantly, because a spinout that discovers a blocking patent after raising money may have to license it on unfavourable terms, redesign the product or stop. A full freedom-to-operate opinion is expensive and is usually commissioned only when a product design is reasonably settled. But an early landscape search, identifying the main patent holders in the field and any obvious blocking rights, is cheap relative to the cost of discovering a problem late, and it helps the inventor design around obstacles while the technology is still flexible. When patents are the wrong tool Patents are not always the best form of protection. For some inventions, especially manufacturing processes that cannot be discovered by examining the final product, keeping the method secret may protect it longer and more cheaply than patenting it, since a trade secret does not expire as long as it stays secret, and a patent publishes the method for competitors to study. For software and data-driven methods, patents may be hard to obtain after the US decisions described earlier, and speed, data and know-how may matter more. For some research tools, a university may decide that open licensing or no patent at all will serve the public better and create more value than exclusivity. Most deep-tech ventures end up protected by a combination: core patents on the key composition or device, further patents on improvements and applications, trade secrets around manufacturing, and accumulated know-how in the team. The patent, however, remains central for most of them, because it is the one form of protection that can be transferred cleanly by licence, valued by an investor and enforced against a well-funded competitor. How a single filing grows into a portfolio, and how that portfolio moves from the university into a company, is the subject of the next chapter. Hashtags: #ResearchCommercialization #AcademicInnovation #DeepTechSpinouts #TechnologyTransfer #ValleyOfDeath #AcademicPatents #IntellectualProperty #BayhDoleAct #InventionDisclosure #PatentStrategy #PatentPortfolio #FreedomToOperate #TechnologyLicensing #SpinoutLicensing #TranslationalFunding #NonDilutiveFunding #TechnologyReadinessLevels #ProofOfConcept #DeepTechEntrepreneurship #FounderEquity #VentureCapital #TermSheets #ConflictOfInterest #UniversitySpinouts #FutureOfResearchCommercialization

  • Research Animal Welfare and 3Rs Strategy (Replacement, Reduction, and Refinement)

    Download the Book (PDF): Introduction Every protocol that involves a living animal begins with a question that is easy to skip and hard to answer well: why this animal, why this many, and why this way? The question is not a formality. It sits at the centre of the modern law of animal research in Europe, North America and much of the rest of the world, and it is the practical form of a principle first set out more than sixty years ago. Asked seriously, it improves the lives of the animals involved. Asked seriously, it also improves the science. That second point is the argument of this book. The framework known as the Three Rs, Replacement, Reduction and Refinement, is usually presented as an ethical constraint on research: a set of obligations that scientists owe to animals and that regulators enforce, sometimes against the grain of scientific convenience. That picture is not wrong, but it is incomplete in a way that matters. Properly understood, the Three Rs are a discipline of good experimental practice. An animal that suffers unnecessarily is a source of uncontrolled variation. A study that uses too few animals to detect the effect it is looking for wastes every animal in it. A model that does not resemble the human condition it is meant to represent produces results that will not translate, however many animals are used. A replacement method that captures the relevant human biology more faithfully than a rodent is not a compromise; it is a better instrument. Welfare and validity are not rival goods to be traded off against each other. In most of the decisions that matter, they point the same way. Where the idea came from The Three Rs were formulated by the zoologist William Russell and the microbiologist Rex Burch in The Principles of Humane Experimental Technique, published in 1959 after a study commissioned by the Universities Federation for Animal Welfare. Their book was neither a polemic against animal research nor a defence of it. It was an attempt to apply scientific thinking to the problem of inhumanity in the laboratory: to classify its sources, to identify the techniques that could reduce it, and to argue that the most humane science and the best science tended to coincide. Russell and Burch wrote that "the greatest scientific experiments have always been the most humane and the most aesthetically attractive", and they meant it as an empirical observation rather than a pious hope. For two decades the book was largely neglected. Its revival from the late 1970s onwards, and its absorption into law, guidance and institutional practice from the 1980s, turned three words into an organising principle for a whole sector. Today the Three Rs appear in the text of European Union Directive 2010/63/EU, which obliges member states to ensure that the principles of replacement, reduction and refinement are applied systematically. They are woven through the United States' regulatory system, from the Animal Welfare Act's requirement that investigators consider alternatives to painful procedures to the Guide for the Care and Use of Laboratory Animals, which endorses them explicitly. National centres now exist to advance them, most prominently the UK's National Centre for the Replacement, Refinement and Reduction of Animals in Research (NC3Rs), founded in 2004. Journals ask authors to report their studies according to the ARRIVE guidelines, which exist in large part because poor reporting makes animal studies impossible to evaluate or reproduce. What this book covers The chapters that follow set out what a rigorous, welfare-centred approach to animal research looks like in practice, from the original concepts through the regulatory architecture to the concrete decisions made at the bench, in the animal facility and in the ethics committee. Chapter 1 returns to Russell and Burch to recover what the Three Rs actually meant and why their authors saw humaneness and scientific quality as connected. Several ideas in their book, including the distinction between high-fidelity and discrimination models and the warning against assuming that a more similar animal is always a better model, are more useful today than the slogan that survives them. Chapter 2 describes the regulatory frameworks within which animal research operates: the European Directive with its harm-benefit analysis and severity classification, the parallel but differently constructed systems of the United States, and the United Kingdom's licensing regime. The purpose is not to turn readers into compliance officers but to show how the law turns an ethical principle into a sequence of decisions, and where those decisions are made. Chapters 3 and 4 take up Replacement. The first surveys the landscape of established non-animal methods, with particular attention to regulatory toxicology, where validated replacements have displaced animal tests entirely for some endpoints. The second turns to the newer human-based technologies that are reshaping biomedical research: organoids grown from stem cells, microfluidic organ-on-chip systems, and computational models that simulate physiology and predict toxicity. It asks what these methods can already do, what they cannot yet do, and how their validity should be judged. Chapters 5 and 6 are about Reduction, which is best understood as the application of sound experimental design and statistics. Chapter 5 addresses bias, randomisation, blinding and transparent reporting, drawing on the evidence that much published animal research has been compromised by weaknesses that are cheap to prevent. Chapter 6 treats sample size in detail: why studies that are too small waste animals as surely as those that are too large, how power calculations work and what they require, and which design strategies allow a given question to be answered with fewer animals. Chapters 7 and 8 address Refinement. Chapter 7 concerns the recognition and prevention of pain, the planning of anaesthesia and analgesia as a matter of institutional policy and veterinary partnership, and the evidence that withholding pain relief is far less often scientifically justified than was once assumed. Chapter 8 deals with humane endpoints, severity assessment and the animal's whole lifetime experience, from breeding and housing to handling and the manner of its death. Throughout these chapters, pain management is discussed at the level of principle, planning and professional responsibility. The specific choice of drugs, doses and schedules is a clinical matter for the attending veterinarian working with the research team and the relevant formulary, and it is not the business of a book like this to substitute for that consultation. Chapter 9 turns to the institution: the ethics committees and welfare bodies that review protocols, the training and competence of the people who handle animals, the "culture of care" that determines whether the rules are followed in spirit as well as letter, and the role of openness, retrospective review and systematic evidence synthesis in making the whole enterprise accountable. The conclusion draws these threads together and considers what remains unsettled, including the fast-moving policy shifts toward human-based methods in regulatory science and the unresolved questions about how new approaches should be validated. Who it is for The book is written for anyone who designs, reviews, supports or oversees research involving animals: early-career scientists writing their first protocols, established investigators who want to see how the field's expectations have moved, members of ethics committees, including the lay and non-scientific members who bring an essential outside perspective, animal technologists and care staff, veterinarians new to laboratory animal medicine, research managers, and funders. It should also be useful to students and to members of the public who want to understand how animal research is regulated and how the people who conduct it think about their responsibilities. It assumes no specialist knowledge of statistics, pharmacology or law. Where technical ideas matter, such as statistical power or the validation of an alternative method, they are explained in plain terms. The regulatory detail is accurate to the published frameworks, but readers should always consult the current text of the rules that apply to them and the guidance of their own institution, because both change. A note on stance This is not a book that argues for or against the use of animals in science as such. Reasonable people disagree about that question, and the disagreement is not going to be settled by a better protocol. What can be said with confidence is that animals are currently used in large numbers in research, that the law in most jurisdictions permits this only under conditions, and that within those conditions there is almost always room to do better: to replace an animal study with a method that answers the question more directly, to design an experiment that needs fewer animals and gives a clearer answer, to spare an animal pain that serves no purpose. The people best placed to make those improvements are the people who do the work. This book is written for them, in the conviction that compassion and rigour are not separate virtues in the laboratory but the same habit of taking the details seriously. Chapter 1: What Russell and Burch Actually Argued The Three Rs are among the most widely cited ideas in the life sciences and among the least read. Many researchers who can recite the three words have never opened the book in which they were introduced, and much of what that book contains has been lost in the translation from argument to slogan. Recovering it is worth the effort, because the original is more demanding, more scientific and more practical than the version that circulates in training slides. The study and its setting In 1954 the Universities Federation for Animal Welfare, a British charity founded by Charles Hume with a distinctive commitment to working with scientists rather than against them, commissioned a systematic study of humane technique in laboratory experimentation. The task went to William Russell, a zoologist with broad interests in psychology and the classics, assisted by Rex Burch, a microbiologist. Their report, The Principles of Humane Experimental Technique, appeared in 1959. The period matters. Animal use in British and American laboratories was expanding rapidly after the Second World War, driven by the growth of pharmacology, the demands of vaccine production and biological standardisation, and a new scale of government-funded biomedical research. The regulatory framework in Britain was still the Cruelty to Animals Act of 1876, which licensed individual experimenters but said little about how experiments should be designed. The United States had no federal law governing laboratory animal welfare at all; the Animal Welfare Act would not arrive until 1966. Into this setting Russell and Burch brought something unusual: an attempt to treat inhumanity itself as a phenomenon that could be studied, classified and reduced by scientific means. The definitions Russell and Burch defined their three principles with some care, and the definitions are worth stating accurately. Replacement meant any scientific method that uses non-sentient material in place of conscious living vertebrates. They included under this heading the use of higher plants, microorganisms, tissue and cell cultures, and what they called "physico-chemical" techniques, and they distinguished between absolute replacement, in which no animal is used at any stage, and relative replacement, in which animals are still required, for instance as a source of tissue, but are not subjected to any procedure that could cause distress while alive. Reduction meant lowering the number of animals needed to obtain information of a given amount and precision. The qualification is essential and frequently forgotten. Reduction is not simply using fewer animals; it is using fewer animals for the same quality of answer. A study that halves its animal numbers but can no longer detect the effect it was designed to find has not achieved reduction. It has converted a useful experiment into a wasteful one. Refinement meant any decrease in the incidence or severity of inhumane procedures applied to those animals that still have to be used. Russell and Burch were explicit that refinement applies to everything that happens to the animal, not only to the experimental procedure, and that its aim is to reduce the total amount of distress the animal experiences. Two features of this scheme deserve emphasis. First, the definitions are ordered by a logic of priority. Replacement comes first because, where it is possible, it removes the problem entirely. Reduction and refinement apply to the animals that still have to be used. Second, the three principles are not independent. Reduction can conflict with refinement, for example when fewer animals are each subjected to more procedures, and later writers have spent considerable effort on how such conflicts should be resolved. Russell and Burch themselves were clear that the goal was to minimise the total inhumanity of the enterprise, not to maximise any one of the three Rs at the others' expense. Humane and inhumane The book's central category was not "the Three Rs" but inhumanity, which Russell and Burch analysed with some subtlety. They distinguished direct inhumanity, the distress inflicted as an unavoidable consequence of a procedure, from contingent inhumanity, which arises as an incidental and inessential by-product of the way the work is done: poor husbandry, rough handling, inadequate anaesthesia, avoidable infection, neglect of the animal's needs outside the experimental procedure itself. Much of the suffering in laboratories, they argued, was contingent. It served no scientific purpose, and in many cases it actively undermined the science by introducing stress and disease into the very animals whose physiology was being measured. This distinction remains one of the most practically useful ideas in the field. When a protocol is reviewed today, the question of direct harm, what the procedure itself does to the animal, is usually well considered, because it is written down and scrutinised. Contingent harm is harder to see. It lives in the details: the temperature of the animal room, the method of picking a mouse up, whether a post-operative animal can reach its food and water, whether the person checking the animals at the weekend knows what signs to look for. Refinement in the fullest sense is largely the systematic elimination of contingent inhumanity, and it is where the most improvement is usually available at the least scientific cost. Fidelity and discrimination One of the more sophisticated sections of The Principles of Humane Experimental Technique concerns the nature of models. Russell and Burch distinguished between two properties a model might have. A high-fidelity model resembles the system being modelled in many respects; the classic assumption is that an animal more closely related to humans, or more similar physiologically, is a better model. A discrimination model, by contrast, is one that reproduces the specific property under investigation, even if it differs from the target in almost every other respect. They argued that researchers were prone to what they called the "high-fidelity fallacy": the assumption that a model's overall similarity to humans guarantees its relevance to a particular question. In fact, what matters is whether the model behaves like the target system with respect to the mechanism being studied. A cell line expressing the human version of a receptor may be a far better model for a drug's action at that receptor than a whole rodent whose receptor differs in its binding properties. Conversely, an animal that shares many features of human physiology may diverge from humans on exactly the pathway that matters, and its apparent fidelity then becomes a source of false confidence. This argument is the intellectual foundation for replacement as a scientific strategy rather than a moral concession. It explains why, in some fields, human-based in vitro and computational methods have outperformed animal models on specific predictive tasks, and why the value of an animal model has to be established for each question rather than assumed. The failure of many treatments that worked in animal models of stroke, sepsis, amyotrophic lateral sclerosis and other conditions to translate into benefit in patients is, in part, a demonstration of the high-fidelity fallacy at scale. Researchers treated whole-animal models as faithful representations of human disease when the relevant mechanisms differed in ways that mattered. Why the book was neglected, and why it returned For most of the 1960s and 1970s, the Three Rs had little visible influence. Russell himself later suggested various reasons: the book was long and dense, the scientific community of the time was not receptive, and the anti-vivisection movement had little interest in a framework that accepted animal research as legitimate while seeking to improve it. The period's expansion of research also made reflection on its methods seem less urgent than getting on with the work. The revival came from several directions. In the late 1970s and 1980s, public controversy over animal testing, particularly of cosmetics and household products, created pressure for alternatives. Organisations such as the Fund for the Replacement of Animals in Medical Experiments in the UK, founded in 1969, and the Johns Hopkins Center for Alternatives to Animal Testing, established in 1981, began to champion the approach. Regulatory reforms, notably the United States' 1985 amendments to the Animal Welfare Act and the United Kingdom's Animals (Scientific Procedures) Act 1986, incorporated elements of the Three Rs into law. By the 1990s the framework had become the common language of research ethics committees, regulators and welfare scientists across much of the world. Interpretive drift Popularity came at a price. In their 2015 paper in the Journal of the American Association for Laboratory Animal Science, Jerrold Tannenbaum and B. Taylor Bennett examined how the definitions of the Three Rs had shifted since 1959 and argued that inconsistency in how they are defined and applied creates real confusion in practice. Some institutions define refinement narrowly, as the reduction of pain and distress during procedures; others extend it to the enhancement of positive welfare throughout an animal's life. Some treat reduction as any decrease in numbers; others insist, as Russell and Burch did, on the link to information of a given precision. Some treat replacement as applying only to methods that eliminate animal use entirely; others count the use of less sentient animals, such as invertebrates or early developmental stages, as a form of replacement. These are not merely semantic disputes. How the terms are defined determines what counts as compliance, what research gets funded as "3Rs research", and what a protocol reviewer is entitled to ask for. The modern trend, reflected in guidance from the NC3Rs and in the wording of the European Directive, is towards broader and more ambitious definitions. Refinement is increasingly understood to encompass not only the minimisation of pain, suffering, distress and lasting harm but also the promotion of positive welfare: the opportunity for animals to perform motivated behaviours, to experience comfort and to have some control over their environment. Replacement is increasingly defined as the use of methods that avoid or replace the use of animals, with the use of invertebrates or immature forms treated as a partial replacement rather than a full one. The Three Rs as a scientific discipline The most important thing to recover from Russell and Burch is their conviction that humane technique and good science are allied. They did not claim that every humane improvement would also improve the science, or that no scientifically valuable experiment would ever require animal suffering. They claimed something more modest and more robust: that inhumanity in the laboratory is, much more often than researchers assume, a symptom of poor technique, and that the effort to reduce it tends to expose and correct scientific weaknesses at the same time. The evidence gathered since 1959 has largely vindicated this view. Stressed animals show altered immune function, endocrine profiles, metabolism and behaviour, all of which can confound experimental outcomes. Animals in pain may eat less, move less and sleep poorly, with downstream effects on almost any physiological measurement. Studies with inadequate sample sizes produce unreliable results that cannot be replicated, so that the animals used in them were used for nothing. Poorly designed experiments that fail to randomise or blind are vulnerable to bias that inflates effect sizes, leading whole fields down paths that later collapse. In each case, the welfare problem and the scientific problem are the same problem seen from different sides. This convergence should not be overstated. There are genuine cases in which welfare and scientific objectives conflict: an analgesic that interferes with the inflammatory process under study, an endpoint that must be observed at a stage of disease that involves suffering, a housing refinement that introduces variation into a carefully controlled environment. These conflicts are real, and the later chapters of this book address how they should be handled. But they are fewer than is often supposed, and the default assumption, that welfare improvements cost scientific quality, is usually wrong. The more productive starting point, and the one Russell and Burch proposed, is that each welfare problem in a protocol should be examined as a possible scientific problem too. When the Rs pull in different directions Because the three principles are applied to the same animals and the same experiments, they sometimes compete, and a mature application of the framework requires judgement about how to weigh them. Three kinds of tension recur. The first is between reduction and refinement at the level of the individual animal. Longitudinal designs, in which the same animals are measured repeatedly over time, can dramatically reduce the number of animals needed, because each animal serves as its own control and between-animal variation is removed from the comparison. But repeated measurement may mean repeated anaesthesia, repeated blood sampling, repeated handling or repeated imaging, each of which carries some welfare cost. The reuse of animals across separate procedures raises the same issue in a sharper form. The European Directive addresses this explicitly: an animal that has already been used in one or more procedures may be reused only if the actual severity of the previous procedures was mild or moderate, a veterinarian has advised in favour, its general state of health and well-being has been fully restored, and the further procedure is classified as no more than moderate, subject to limited exceptions. The underlying principle is that reduction in numbers cannot be purchased with an unacceptable increase in the cumulative burden borne by the animals that remain. The second tension is between replacement and the other two Rs at the level of a research programme. Developing and validating a replacement method may itself require animal studies, for instance to generate the reference data against which the new method is compared. In the short term this can increase animal use; the case for it rests on the long-term reductions that a validated alternative will deliver. Deciding when such investment is worthwhile is a strategic judgement, and it is one reason why funders such as the NC3Rs support replacement research directly rather than leaving it to individual laboratories. The third tension concerns species choice. A study might be performed on fewer animals of a species with greater capacity for suffering, or on more animals of a species with less. Russell and Burch did not provide a formula for this, and none has been generally accepted since. Modern practice tends to prefer the species of lowest neurophysiological sensitivity that can answer the question validly, which is one reason why zebrafish larvae before the age of independent feeding, Drosophila and nematodes have become attractive for early-stage work. But validity comes first: using a less sentient organism that cannot model the relevant biology simply wastes those organisms and delays the question. These tensions are not failures of the framework. They are the reason it requires trained human judgement to apply, and the reason the law places that judgement in the hands of ethical review bodies rather than reducing it to a checklist. Beyond the original three The Three Rs sit within a wider body of thinking about animal welfare that has developed alongside them. The Brambell Committee, reporting to the UK Parliament in 1965 on the welfare of farmed animals, gave rise to what became known as the Five Freedoms: freedom from hunger and thirst, from discomfort, from pain, injury and disease, from fear and distress, and freedom to express normal behaviour. Although developed for agriculture, the Five Freedoms shaped thinking about laboratory housing and husbandry for decades. More recently, the Five Domains model developed by David Mellor and colleagues has moved the emphasis from the absence of negative states towards the presence of positive ones, assessing nutrition, physical environment, health and behavioural interactions as inputs to the animal's overall mental state. This shift, from avoiding harm towards promoting a life worth living, has been absorbed into modern interpretations of refinement. It matters for laboratory animals because many of them spend the great majority of their lives not undergoing procedures but living in cages. What happens to them in those hours, how much space and complexity they have, whether they can nest, burrow, hide, climb or interact with companions, is as much a part of their welfare as the procedure itself, and often a larger part. From principle to practice What turns this conviction into practice is a set of questions that can be asked of any proposed study. Could the question be answered without using live animals, or with fewer, or with less harm? If animals are necessary, is the chosen species and model the one that best represents the mechanism of interest, rather than the one that is most familiar or convenient? Is the experiment designed so that its results will be reliable, with appropriate controls, randomisation, blinding and sample size? Has every source of avoidable harm, direct and contingent, been identified and addressed? Are the endpoints set at the earliest point that answers the scientific question? Will the results be reported in enough detail for others to evaluate and build on them, so that the animals' use is not wasted by an irreproducible paper? None of these questions is new, and none of them is sufficient on its own. Together they constitute the working form of the Three Rs, and the rest of this book is an elaboration of how to answer them well. The regulatory frameworks described in the next chapter are, in large part, institutional machinery for making sure that these questions are asked, by someone, before any animal is used. Chapter 2: The Regulatory Architecture The Three Rs became powerful when they became law. A principle that scientists were invited to consider turned into a set of obligations that must be satisfied before a project can begin, reviewed while it runs and reported on after it ends. The legal frameworks differ considerably between jurisdictions, in what they cover, who enforces them and how much discretion they leave to institutions. But all of the major systems share a basic structure: an animal may be used in science only when a responsible body has concluded that the use is justified, that its harms have been minimised and that the people doing the work are competent to do it. This chapter sets out the main frameworks in outline. It is not a substitute for the text of the regulations or for institutional guidance, both of which change and both of which govern in detail. Its purpose is to show how each system translates the Three Rs into decisions, and where in the life of a project those decisions fall. The European Union: Directive 2010/63/EU The most comprehensive codification of the Three Rs in law is Directive 2010/63/EU on the protection of animals used for scientific purposes, adopted in September 2010 and applicable in member states from January 2013. It replaced a 1986 directive that had become outdated and inconsistently applied, and it was explicitly designed to raise and harmonise standards across the Union. The Directive states in its recitals that its ultimate goal is the full replacement of procedures on live animals for scientific and educational purposes as soon as it is scientifically possible to do so. Until then, it seeks to ensure the highest possible standard of protection for the animals that are used. Its scope is broad. It covers live non-human vertebrates, including independently feeding larval forms and foetal forms of mammals from the last third of their normal development, and live cephalopods such as octopus and squid. The inclusion of cephalopods, on the basis of evidence of their capacity to experience pain, suffering, distress and lasting harm, was a notable extension compared with most other jurisdictions. The Three Rs are written into the Directive as legal requirements rather than aspirations. Article 4 requires member states to ensure that, wherever possible, a scientifically satisfactory method or testing strategy not entailing the use of live animals is used instead of a procedure; that the number of animals used is reduced to a minimum without compromising the objectives of the project; and that breeding, accommodation and care, and the methods used in procedures, are refined so as to eliminate or reduce to a minimum any possible pain, suffering, distress or lasting harm. Article 13 adds that a procedure must not be carried out if another method not entailing the use of a live animal, and recognised under Union legislation, exists; and that in choosing between procedures, those that use the minimum number of animals, involve animals with the lowest capacity to experience pain, suffering, distress or lasting harm, cause the least of those harms and are most likely to provide satisfactory results should be selected. Death as an endpoint is to be avoided as far as possible and replaced by early and humane endpoints. Project evaluation and the harm-benefit analysis The mechanism that brings these requirements to bear on individual research is project authorisation. Under Article 36, no project may be carried out without prior authorisation by the competent authority, and authorisation depends on a favourable project evaluation. Article 38 specifies what the evaluation must include: verification that the project is justified from a scientific or educational point of view or required by law; verification that its purposes justify the use of animals; verification that it is designed so as to enable procedures to be carried out in the most humane and environmentally sensitive manner possible; an assessment of the project's compliance with the Three Rs; a classification of the severity of the procedures; and a harm-benefit analysis, to assess whether the harm to the animals in terms of suffering, pain and distress is justified by the expected outcome, taking into account ethical considerations, and may ultimately benefit human beings, animals or the environment. The harm-benefit analysis is the heart of the system and its most contested element. It requires reviewers to weigh incommensurable quantities: the suffering of particular animals against the prospect of knowledge or benefit that is uncertain, delayed and often indirect. No formula makes this weighing mechanical, and the Directive does not attempt to provide one. What it does require is that the weighing be done explicitly, by people with appropriate expertise, before the project begins. The discipline of writing down what harms are expected, what benefits are claimed and how likely those benefits are is itself a significant safeguard. It forces the applicant to confront the question of whether the study is capable of delivering what it promises, which is often where weak designs are exposed. Severity classification Article 15 requires every procedure to be classified, prospectively, into one of four categories set out in Annex VIII: non-recovery, mild, moderate or severe. The classification is based on the degree of pain, suffering, distress or lasting harm expected to be experienced by an individual animal during the course of the procedure. Annex VIII provides criteria and examples for each category, and the European Commission has published further guidance in the form of working documents on the severity assessment framework. The categories matter in several ways. They inform the harm side of the harm-benefit analysis. They determine whether a project requires retrospective assessment. They constrain reuse of animals, as described in the previous chapter. And they impose an upper limit: Article 15 provides that member states must ensure that a procedure is not performed if it involves severe pain, suffering or distress that is likely to be long-lasting and cannot be ameliorated. A safeguard clause allows a member state, in exceptional and scientifically justified circumstances, to permit such a procedure provisionally, subject to notification and review at Union level. The practical effect is that there is a ceiling on permitted suffering, not merely a requirement to justify it. The Directive also requires that the actual severity experienced by each animal be recorded and reported, not just the prospective classification. Annual statistical reports from member states therefore include the actual severity of procedures, which allows comparison between what was predicted and what happened. This retrospective reporting has had a useful disciplinary effect: it creates a record against which predictions can be checked, and it exposes procedures whose real welfare cost was higher than anticipated. Retrospective assessment, welfare bodies and transparency Article 39 requires retrospective assessment of projects that use non-human primates or involve procedures classified as severe, and allows competent authorities to require it for others. The assessment evaluates whether the objectives of the project were achieved, the harm actually inflicted on animals, including the numbers and species used and the severity of the procedures, and any elements that may contribute to further implementation of the Three Rs. Its purpose is learning: to find out whether the harm-benefit judgement made at the outset was borne out, and to feed what was learned into future projects. Each breeder, supplier and user must establish an animal-welfare body under Article 26, including at least the person responsible for welfare and care and, in the case of a user, a scientific member. The body advises staff on welfare, advises on the application of the Three Rs, establishes and reviews internal processes for monitoring and reporting, follows the development and outcome of projects, and advises on rehoming schemes. Article 25 requires each establishment to have a designated veterinarian with expertise in laboratory animal medicine. Transparency is addressed through Article 43, which requires non-technical project summaries to be published, giving the objectives of the project, predicted harms and benefits, the number and types of animals, and a demonstration of compliance with the Three Rs. The summaries are written for a general audience and are now collected in a publicly searchable European database. The United States: two overlapping systems The United States regulates animal research through two frameworks that overlap but differ in scope, legal basis and enforcement. Understanding their relationship is essential for anyone working in or with American institutions. The Animal Welfare Act The Animal Welfare Act, originally passed in 1966 as the Laboratory Animal Welfare Act and amended several times since, is a federal statute enforced by the Animal and Plant Health Inspection Service of the US Department of Agriculture. Its regulations are set out in Title 9 of the Code of Federal Regulations. The 1985 amendments, known as the Improved Standards for Laboratory Animals Act, were particularly significant for research. They required research facilities to establish Institutional Animal Care and Use Committees, required investigators to consider alternatives to procedures that may cause more than momentary or slight pain or distress, required consultation with a veterinarian in planning such procedures, required the use of anaesthetics, analgesics and tranquillisers unless withholding them is scientifically justified, and introduced requirements for the exercise of dogs and the psychological well-being of non-human primates. The Act's coverage is limited in a way that surprises many people outside the field. Its definition of "animal" excludes birds, rats of the genus Rattus and mice of the genus Mus bred for use in research, as well as cold-blooded animals and farm animals used for agricultural research. Since rats and mice bred for research make up the great majority of animals used in American laboratories, most research animals fall outside the Act's direct protection. The species it does cover include dogs, cats, non-human primates, rabbits, guinea pigs and hamsters, among others. The USDA requires registered facilities to report annually the numbers of covered animals used, classified by pain category, including whether procedures involving pain or distress were conducted with or without appropriate relief. The PHS Policy and the Guide The gap left by the Act's exclusions is largely filled, for publicly funded research, by the Public Health Service Policy on Humane Care and Use of Laboratory Animals, administered by the Office of Laboratory Animal Welfare at the National Institutes of Health. The PHS Policy applies to all live vertebrate animals used in research, research training or testing conducted or supported by PHS agencies, including the NIH. Institutions receiving such funding must file an Animal Welfare Assurance with OLAW, describing their programme for animal care and use and committing to comply with the Policy. The PHS Policy requires institutions to use the Guide for the Care and Use of Laboratory Animals, published by the National Research Council, as the basis for their programmes. The current eighth edition was published in 2011. The Guide is a substantial document covering the institutional programme, the animal environment, housing and management, veterinary care and the physical plant. It endorses the Three Rs explicitly and adopts a largely performance-based approach, specifying the outcomes that must be achieved while allowing institutions flexibility in how they achieve them. The Policy also incorporates the US Government Principles for the Utilization and Care of Vertebrate Animals Used in Testing, Research, and Training, a short set of principles issued in 1985 which, among other things, state that procedures with animals should avoid or minimise discomfort, distress and pain consistent with sound scientific practices, and that investigators should consider appropriate alternatives such as mathematical models, computer simulation and in vitro biological systems. Many institutions additionally seek voluntary accreditation from AAALAC International, which assesses programmes against the Guide and other standards and is widely regarded as a mark of quality in the United States and internationally. The IACUC Under both frameworks, the central decision-making body is the Institutional Animal Care and Use Committee. The IACUC must include at least a veterinarian with programme responsibility, a practising scientist experienced in animal research, a member whose primary concerns are in a non-scientific area, and a member not affiliated with the institution. It reviews and approves, requires modifications to, or withholds approval of proposed animal activities; conducts semi-annual reviews of the institutional programme and inspections of facilities; investigates concerns; and has the authority to suspend activities that are not being conducted in accordance with approved protocols. The IACUC's review criteria mirror the Three Rs, though they are expressed differently from the European framework. The committee must determine, among other things, that procedures avoid or minimise discomfort, distress and pain; that the investigator has considered alternatives to painful procedures and has provided a written narrative of the methods and sources used to determine that alternatives were not available; that the activities do not unnecessarily duplicate previous experiments; that appropriate sedation, analgesia or anaesthesia will be used; and that animals that would otherwise experience severe or chronic pain or distress that cannot be relieved will be painlessly euthanised at the end of the procedure or, if appropriate, during it. The requirement to justify the number of animals requested is standard in IACUC practice and embedded in the PHS Policy's requirement that the protocol describe the rationale for the number of animals. What the American system does not require, in contrast to the European one, is an explicit harm-benefit analysis with prospective severity classification in the Directive's sense. The IACUC considers justification and minimisation of harm, but its mandate is framed more around the adequacy of the welfare provisions for a given study than around a formal weighing of the study's value against its costs. How much this difference matters in practice is debated; many IACUCs engage in something close to harm-benefit reasoning, while others focus narrowly on welfare procedures and leave scientific merit to peer review. The United Kingdom The United Kingdom has regulated animal research by statute since 1876, and its current framework, the Animals (Scientific Procedures) Act 1986, was amended in 2012 to transpose the European Directive. The amended Act was retained after the UK left the European Union, so its substantive requirements remain closely aligned with the Directive, including harm-benefit analysis, severity classification, retrospective assessment and the protection of cephalopods. The British system is distinctive in its three-tier licensing structure, administered by the Home Office. An establishment licence covers the premises. A project licence authorises a programme of work and is granted only after a harm-benefit assessment. A personal licence authorises an individual to carry out regulated procedures, and is granted only to people who have completed accredited training. Each establishment must appoint named persons with defined responsibilities: a Named Veterinary Surgeon, a Named Animal Care and Welfare Officer, a Named Information Officer and a Named Training and Competency Officer, together with an Animal Welfare and Ethical Review Body that fulfils the role of the Directive's animal-welfare body and more. The Home Office's inspectorate conducts visits and audits. Non-technical summaries of all project licences are published. Table 1 sets the main features of these frameworks side by side. Table 1. Main features of three regulatory frameworks for animal research. Feature EU Directive 2010/63/EU United States (AWA and PHS Policy) United Kingdom (ASPA 1986, amended 2012) Species covered Live non-human vertebrates, some foetal and larval forms, cephalopods AWA: excludes purpose-bred rats, mice and birds; PHS: all live vertebrates in funded work As EU, including cephalopods Review body Competent authority, advised by animal-welfare body IACUC Home Office, advised by AWERB Harm-benefit analysis Explicit legal requirement Not formally required Explicit legal requirement Severity classification Prospective and actual, four categories USDA pain categories for covered species Prospective and actual, four categories Retrospective assessment Required for primates and severe procedures Not required Required on similar basis Public summaries Non-technical summaries published Not required Non-technical summaries published Recent shifts in policy toward human-based methods The regulatory landscape is moving, and in the past few years the movement has been noticeably towards encouraging or requiring non-animal methods where they are adequate. In the European Union, the prohibition on animal testing of cosmetic ingredients and finished products, fully in force for marketing purposes since 2013, created one of the first sectors in which animal testing was effectively removed, and it drove substantial investment in validated alternatives. In the United States, the FDA Modernization Act 2.0, signed into law in December 2022, removed the statutory language that had been read as requiring animal testing before human trials of new drugs, making clear that sponsors can use suitable non-animal methods, including cell-based assays, microphysiological systems and computer models, to support applications. In 2025 the FDA announced a plan to reduce and eventually replace animal testing requirements in some areas of preclinical safety assessment, beginning with monoclonal antibodies, and the NIH announced initiatives to prioritise human-based research approaches. The details and pace of these changes continue to develop, and anyone relying on them should consult the current guidance of the agencies concerned. These shifts do not remove the need for animal studies in many areas, and regulators continue to require them where no adequate alternative exists. What they change is the burden of justification. It is increasingly the animal study, rather than the alternative, that must explain why it is the right method for the question. That shift, from treating animal data as the default to treating it as one option among several, each to be justified for a given context, is the regulatory expression of the argument Russell and Burch made about models. What the law cannot do Legal frameworks set floors, not ceilings. They ensure that certain questions are asked and certain standards met, and they provide mechanisms for enforcement when things go wrong. They cannot ensure that the questions are asked thoughtfully or that the standards are embraced rather than merely satisfied. An IACUC or animal-welfare body can approve a protocol that complies with every requirement and still fails to use the best available design, the most suitable model or the most refined technique, simply because nobody involved knew of a better option or thought to look for one. This is why the Three Rs cannot be delegated entirely to compliance. The regulatory frameworks create the occasions on which improvements can be made, at the time of project design, at review, during monitoring and at retrospective assessment. Whether improvements are actually made depends on the knowledge and commitment of the people involved. The remaining chapters of this book are about that knowledge: what replacement, reduction and refinement look like in concrete terms, and how to recognise the opportunities for each. Chapter 3: Replacement in Practice Replacement is the first of the Three Rs and the most ambitious, because where it succeeds it removes the welfare problem altogether rather than mitigating it. It is also the one most often misunderstood. It is sometimes presented as a single technology waiting to be invented, a machine that will one day make animal research unnecessary. In reality it is a large and heterogeneous collection of methods, some decades old and some very new, each suited to particular questions and unsuited to others. The practical skill of replacement lies less in knowing about any one method than in knowing how to break a research question into parts, some of which can be answered without animals, and in recognising when a non-animal approach answers a question more directly than an animal study could. Kinds of replacement It helps to distinguish several kinds of replacement, because they differ in what they achieve and in how easily they can be adopted. Full replacement avoids the use of any animal or animal-derived material. It includes research on human volunteers, human tissues and cells, established cell lines, computational models and, in some classifications, non-biological physical or chemical methods. Examples range from the use of human-derived cell lines in drug screening to studies in human participants using non-invasive imaging, microdosing or carefully controlled challenge models. Partial replacement uses animals, or animal-derived material, in ways that are considered not to cause suffering, or uses organisms thought to have lower capacity for suffering. It includes the use of primary cells or tissues taken from animals killed humanely without any prior procedure, the use of invertebrates such as Drosophila melanogaster or Caenorhabditis elegans, and the use of immature forms of vertebrates, such as zebrafish embryos before the stage at which they feed independently, which are not protected under the European Directive. Whether partial replacement counts as replacement at all is a matter of definitional dispute. The NC3Rs includes it; some others prefer to treat it as a form of refinement. Whatever the label, it is often a valuable step. Russell and Burch's distinction between absolute and relative replacement captures a similar idea. What matters in practice is that the method chosen answers the scientific question at least as well as the animal study it replaces, and preferably better. Human tissues and human volunteers The most direct form of replacement uses human material or human subjects. Human tissue obtained from surgery, biopsies or post-mortem donation, with appropriate consent and ethical approval, allows researchers to study human physiology and pathology directly rather than through an animal proxy. Tissue banks and biobanks have made such material more widely available. Precision-cut tissue slices from human organs can be maintained for days, preserving cellular architecture and some functional properties, and are used in pharmacology, toxicology and disease research. Studies in human volunteers have also expanded as safer and more sensitive methods have developed. Non-invasive imaging, including functional magnetic resonance imaging, positron emission tomography and advanced ultrasound, allows the study of human physiology in health and disease without harm. Microdosing studies, in which volunteers receive very small, sub-pharmacological doses of a drug, combined with highly sensitive analytical techniques such as accelerator mass spectrometry, can provide early information on how a compound is absorbed, distributed and eliminated in humans. Controlled human infection models, in which carefully selected volunteers are deliberately exposed to well-characterised pathogens under close medical supervision, have been used to study malaria, influenza and other infections and to test vaccines, sometimes answering questions that animal models could not because of differences in host-pathogen biology. Such studies are governed by the ethics of human research, not animal research, and require their own careful justification, but where they are appropriate they can offer data of a directness that no animal model can match. Cell culture and its limits Cell culture is the workhorse of non-animal biology. Established cell lines, primary cells and, increasingly, cells derived from human induced pluripotent stem cells are used in virtually every area of biomedical research. The discovery by Shinya Yamanaka and colleagues, published in 2006 for mouse cells and in 2007 for human cells, that adult somatic cells could be reprogrammed into pluripotent stem cells by the introduction of a small number of transcription factors transformed the field. It made it possible to generate human cells of many types, including neurons, cardiomyocytes and hepatocytes, from individual donors, including patients with specific genetic diseases, without the ethical constraints associated with embryonic stem cells. Conventional two-dimensional cell culture has well-known limitations. Cells grown as monolayers on rigid plastic surfaces lose many of the properties they have in tissues: their shape, their polarity, their interactions with neighbouring cells and the extracellular matrix, and often their differentiated functions. Hepatocytes in standard culture, for instance, lose much of their drug-metabolising capacity within days. Immortalised cell lines, adapted over many generations to grow in culture, may have accumulated genetic and phenotypic changes that make them poor representatives of the tissue they came from. Misidentification and cross-contamination of cell lines have been documented problems for decades; HeLa cells in particular have contaminated many other lines. Authentication of cell lines, now expected by many journals and funders, is a basic requirement of reliable in vitro research. A further issue, sometimes overlooked in discussions of replacement, is the use of animal-derived materials in cell culture itself. Foetal bovine serum, still widely used as a culture supplement, is harvested from bovine foetuses at slaughter, raises its own welfare concerns and introduces batch-to-batch variability that undermines reproducibility. Many antibodies used in research are produced in animals, some by methods involving significant suffering; the production of monoclonal antibodies by the ascites method in mice has been largely abandoned in Europe in favour of in vitro production, and a 2020 recommendation by the EU Reference Laboratory for alternatives to animal testing (EURL ECVAM) argued that animal-derived antibodies should be replaced by non-animal-derived affinity reagents where possible. Moving to chemically defined, serum-free media and recombinant reagents is therefore both a replacement measure and a reproducibility measure. Regulatory toxicology: where replacement has advanced furthest The clearest successes of replacement have come in regulatory toxicology, the testing of chemicals, pharmaceuticals, cosmetics and other products for safety under legal requirements. This is partly because regulatory tests are standardised, so that a validated alternative can replace a specific test outright, and partly because political and legal pressure, particularly the European cosmetics testing ban and the REACH chemicals regulation, created strong incentives for alternatives. Validation is the process by which the reliability and relevance of a new method for a defined purpose are established. In Europe it is coordinated by EURL ECVAM, part of the European Commission's Joint Research Centre; in the United States, by the Interagency Coordinating Committee on the Validation of Alternative Methods (ICCVAM), supported by the National Toxicology Program's interagency centre. Internationally, the Organisation for Economic Co-operation and Development develops Test Guidelines that, once adopted, are accepted for regulatory purposes across member countries under the system of Mutual Acceptance of Data. This international acceptance is crucial: a method validated in one country but not accepted elsewhere will not displace animal tests for multinational companies. Several endpoints illustrate how replacement has proceeded. Skin irritation, historically assessed by applying substances to the shaved skin of rabbits, can now be assessed using reconstructed human epidermis models, three-dimensional cultures of human keratinocytes that form a multilayered, differentiated epidermis; OECD Test Guideline 439 describes the approach. Skin corrosion and serious eye damage and eye irritation have similar in vitro and ex vivo methods accepted, reducing reliance on the rabbit tests that were among the most publicly criticised procedures in toxicology. Skin sensitisation offers a particularly instructive case. Sensitisation, the process that leads to allergic contact dermatitis, involves a sequence of biological events: the chemical binds to skin proteins, activates keratinocytes, activates dendritic cells, and triggers proliferation of specific T cells. This sequence has been formally described as an Adverse Outcome Pathway, a structured representation of the chain of events from molecular initiating event to adverse outcome. Individual non-animal tests were developed for each of the early key events: a direct peptide reactivity assay for protein binding, keratinocyte reporter assays for keratinocyte activation, and assays measuring markers of dendritic cell activation. No single test captures the whole pathway, but combinations of them, integrated according to fixed rules, can predict sensitisation hazard. In 2021 the OECD adopted Guideline 497 on defined approaches for skin sensitisation, the first OECD guideline to describe defined combinations of non-animal methods, with fixed data interpretation procedures, as a replacement for animal tests for this endpoint. The guideline reported that the defined approaches performed at least as well as the mouse local lymph node assay when both were compared with human data. This example carries a general lesson. Replacement often does not work by substituting one non-animal test for one animal test. It works by understanding the biology well enough to break the endpoint into mechanistic components, testing each component in the system best suited to it, and integrating the results. The Adverse Outcome Pathway framework, developed under the OECD's auspices, provides a common language for doing this across many toxicological endpoints. Quality control of biological products A second area where replacement has made steady progress, with less public attention, is the testing of biological medicines and vaccines before batches are released for use. Many such products have historically required animal tests for each batch, because their potency or safety could not be adequately characterised by physical or chemical means alone. Because these tests are repeated for every batch of every product, they have accounted for large numbers of animals, and because some of them used death or severe illness as the measured outcome, they have been among the most harmful procedures in routine use. Three examples show how this is changing. Pyrogen testing, which checks injectable products for contamination by substances that cause fever, was for decades performed by injecting the product into rabbits and measuring their body temperature. The bacterial endotoxin test using lysate from the blood of horseshoe crabs replaced much of this, though it raised its own concerns about the capture and bleeding of the crabs; recombinant versions of the key clotting factor now offer an animal-free option. The monocyte activation test, which uses human blood cells to detect the release of fever-inducing signalling molecules, detects both endotoxin and non-endotoxin pyrogens, and the European Pharmacopoeia Commission decided in 2021 to phase the rabbit test out of its texts. Botulinum toxin products, used both medically and cosmetically, were long tested for potency by a lethality assay in mice; in 2011 the US Food and Drug Administration approved a cell-based potency assay developed by the manufacturer of one leading product, and other manufacturers have followed. Vaccine batch testing has increasingly moved from animal challenge tests towards a "consistency approach", in which the quality of each batch is assured by demonstrating that it was produced by a validated process and matches well-characterised reference batches on a panel of in vitro measures. These cases share a pattern. The animal test measured a crude, integrated outcome, such as fever or death. The replacement measures something more specific and mechanistically defined, often in human cells, and it is usually more precise, faster and cheaper. The obstacle to replacement was not scientific impossibility but the regulatory inertia of standardised requirements written into pharmacopoeias and product licences, which had to be changed text by text and product by product. Invertebrates and early life stages For questions that require an intact organism, with its interacting tissues, development and behaviour, invertebrates and the earliest life stages of vertebrates offer partial replacement. The fruit fly Drosophila melanogaster and the nematode Caenorhabditis elegans have been central to genetics and developmental biology for a century and half a century respectively, and many of the pathways that govern cell death, ageing, nutrient sensing and neural development were first worked out in them. Their short life cycles, low cost and genetic tractability make them powerful tools for screening genes and compounds before any vertebrate work is contemplated. Zebrafish embryos and larvae occupy an intermediate position. They are vertebrates, with organ systems broadly homologous to those of mammals, and their transparency in early development allows direct observation of organ formation, blood flow and cell behaviour. Under the European Directive, zebrafish are protected only from the stage at which they begin to feed independently, conventionally taken to be around five days after fertilisation at standard rearing temperatures, so studies confined to earlier stages fall outside the regulatory framework. This makes them attractive for screening, toxicology and developmental studies. Their use still requires justification on scientific grounds, and the welfare of the embryos is a matter of ongoing discussion, since the absence of legal protection is not the same as the absence of any capacity to suffer. But as a first step that can narrow the range of questions, compounds or doses that need to be taken into protected animals, they are an important tool. Tiered and integrated strategies Even where no single alternative can yet replace an animal test, non-animal methods can reduce and focus animal use. Integrated Approaches to Testing and Assessment combine existing information, computational predictions, in vitro data and, where still needed, targeted animal studies into a strategy in which each step informs the next. A chemical that is clearly positive in a validated in vitro assay may not need to be tested in animals at all for that endpoint. A chemical that is similar in structure to well-characterised substances may be assessed by read-across, in which data on the analogues are used to predict its properties, with animal testing reserved for cases where the prediction is uncertain. Large-scale programmes such as the United States' Toxicology in the 21st Century collaboration, known as Tox21, and the Environmental Protection Agency's ToxCast programme have screened thousands of chemicals across hundreds of high-throughput in vitro assays, generating data that can be used to prioritise chemicals for further assessment and to build predictive models. These programmes have not replaced animal toxicology wholesale, and their predictive performance for complex, systemic endpoints remains an active area of research, but they have shifted the field toward a model in which animal studies are the last step in an evidence-gathering process rather than the first. Finding alternatives before a study begins Both the European and American frameworks require researchers to consider alternatives before using animals, and the American framework requires a written narrative describing the search. In practice, the quality of these searches varies widely. A perfunctory database query with the terms "alternative" and the name of the disease is unlikely to find the relevant literature, because methods papers rarely describe themselves as alternatives. A useful search starts from the specific scientific question, breaks it into components, and asks for each component what methods exist for answering it, in any system. It draws on specialised resources, such as the EURL ECVAM database service on alternative methods, the NC3Rs resource hub, the Johns Hopkins Center for Alternatives to Animal Testing and the AWIC, the Animal Welfare Information Center of the US Department of Agriculture's National Agricultural Library, which provides guidance on conducting alternatives searches. Consultation with information specialists, experienced colleagues and the institution's 3Rs or welfare staff often reveals options that database searches miss. The search should cover all three Rs, not replacement alone. The American requirement refers to alternatives to procedures that may cause more than momentary or slight pain or distress, and "alternatives" here is understood to include refinements and reductions, such as less invasive techniques, better analgesia regimes or more efficient designs, as well as replacement methods. The cultural obstacles The limits of replacement are often attributed to the state of the science, and in many cases correctly. Whole-organism physiology, including the interactions between organ systems, the immune response, behaviour and development over a lifetime, remains beyond the reach of any current in vitro system, and questions that depend on these properties may still require animals. But there are also obstacles that have little to do with scientific capability. One is familiarity. Researchers trained in a particular animal model tend to continue using it, and laboratories organised around animal work have infrastructure, expertise and collaborative networks that make switching costly. Another is the expectation, among reviewers, editors and regulators, that animal data will be provided, sometimes regardless of whether it is the most informative evidence for the question at hand. Researchers report being asked by reviewers to add animal experiments to studies that relied on human cells or computational approaches, not because the animal data would answer a question the human data could not, but because animal validation was seen as a necessary mark of seriousness. A third obstacle is funding: developing and validating alternatives takes time and money, and the benefits may be diffuse and long-term. Addressing these obstacles is partly a matter of education and partly of institutional incentive. Funders that support the development of alternatives, journals that accept well-validated human-based studies on their own terms, and regulators that accept non-animal data where it is adequate all shift the balance. The recent policy changes described in the previous chapter are, in part, an attempt to remove the institutional presumption in favour of animal data. Whether they succeed will depend on how quickly the scientific community builds confidence in the newer methods, which is the subject of the next chapter. Hashtags: #ResearchAnimalWelfare #ThreeRsStrategy #Replacement #Reduction #Refinement #HumaneExperimentalTechnique #AnimalResearchEthics #AnimalWelfare #BioriskManagement #HarmBenefitAnalysis #SeverityAssessment #HumaneEndpoints #AnimalPainManagement #ExperimentalDesign #SampleSizePlanning #Randomization #Blinding #ARRIVEGuidelines #NonAnimalMethods #Organoids #OrganOnChip #ComputationalModels #IACUC #CultureOfCare #FutureOfAnimalWelfareScience

  • Quantum Computing for Computational Chemists (Algorithms, VQE, and Molecular Ground States)

    Download the Book (PDF): Introduction Every computational chemist knows the moment. The geometry is sensible, the basis set is adequate, the calculation converges, and the answer is wrong. Not wrong by a rounding error, but wrong in a way that matters: the spin state ordering of an iron complex is inverted, a bond dissociation curve bends upward where it should flatten, a catalytic barrier comes out ten kilocalories per mole too low. The culprit, most of the time, is electron correlation, the part of the electronic energy that a single averaged picture of the electrons cannot capture. Chemistry has spent seventy years building methods that recover it, and those methods are extraordinarily good for most molecules. For a stubborn minority they are not, and that minority includes some of the most interesting molecules in biology and industry. Quantum computing arrived in chemistry with a promise aimed directly at that minority. The argument, first sketched by Richard Feynman in 1982 and made concrete for molecules in 2005, goes like this. The exact wavefunction of a molecule lives in a space whose size grows exponentially with the number of electrons and orbitals. A classical computer must store and manipulate that space explicitly, or find a clever way to avoid it. A quantum computer built from qubits lives in the same kind of space by its nature. So a quantum computer should be able to represent a correlated molecular wavefunction that no classical computer can hold, and extract its energy efficiently. The argument is correct as far as it goes. What it leaves out is almost everything that determines whether a chemist will ever benefit from it: how the wavefunction gets prepared in the first place, how many measurements it takes to read out an energy to useful precision, what noise does to the calculation, how much error correction costs, and, above all, how good the classical competition actually is. The last decade has been an education in all of these, and the education has not always been comfortable for the people selling quantum chemistry as the first killer application. This booklet is written for computational chemists who want to understand the field on its merits. It assumes you know what a Slater determinant is, why coupled cluster is the workhorse of molecular energetics, and what an active space means. It does not assume you know anything about qubits, gates, or quantum algorithms, and it builds those ideas from the ground up in the language of electronic structure wherever possible. The argument of this book The central claim is simple to state and takes the whole book to earn. The variational quantum eigensolver, the algorithm that made quantum chemistry the flagship application of early quantum computers, has turned out to be a superb laboratory and a poor engine. It taught the field how to map molecules onto qubits, how to build chemically meaningful circuits, how to measure molecular Hamiltonians, and how to mitigate hardware noise. But its own bottlenecks, which are measurement cost, noise accumulation, and the difficulty of training its circuits, combine with steadily improving classical methods to make it very unlikely that VQE on noisy hardware will ever compute a molecular ground state that a classical computer cannot. If quantum computers deliver a genuine advantage for molecular ground states, the evidence now points to a later and narrower route: error-corrected machines running phase estimation on a specific class of strongly correlated systems, with classical methods doing most of the work around them. That is not a dismissal. A narrower claim that is true is worth more to a working chemist than a broad one that is not. Knowing exactly where the advantage could lie, and why, is what lets you read a press release, a funding call, or a paper and judge what it has actually shown. Why honesty matters here Quantum computing for chemistry has suffered from a particular kind of overstatement. Demonstrations on hydrogen and lithium hydride, molecules that a laptop solves exactly in milliseconds, were routinely presented as steps toward drug discovery. Hardware experiments were described as "beyond classical" and then reproduced on a single classical workstation within weeks. Resource estimates for landmark problems such as the iron-molybdenum cofactor of nitrogenase were quoted without the assumptions that made them meaningful. At the same time, the field has done some genuinely excellent science, and the most rigorous critiques of quantum advantage have come from inside it. The paper that most sharply questioned whether there is exponential quantum advantage in ground-state chemistry was written by researchers who also build quantum algorithms. The work showing that many quantum machine learning speedups could be matched classically, a line of results now called dequantization, came from a researcher, then an undergraduate, who set out to prove the opposite. The tensor-network simulations that brought several hardware "advantage" claims back to classical reach were done by people who respect the experiments enough to take them seriously. The picture as of late 2026 is more interesting than either the hype or the backlash suggests. Error correction has crossed thresholds that were aspirational five years ago. Superconducting processors now suppress logical errors as code size grows. Trapped-ion machines have run phase estimation for molecular hydrogen on error-corrected logical qubits. Hybrid workflows pairing quantum processors with the largest supercomputers in the world have treated iron-sulfur clusters with active spaces beyond exact diagonalization. And in each case the honest reading is the same: real engineering progress, no demonstrated advantage for a chemical question anyone could not already answer classically. How the book is built The chapters move from the chemistry to the machine, then to the algorithm, then to its limits, and finally to what a practitioner should believe and do. Chapter 1 sets out the electron correlation problem as it actually stands, including what the best classical methods can do and where they fail. Without an accurate picture of the competition, no claim of quantum advantage can be evaluated. Chapter 2 shows how a molecular Hamiltonian becomes an operator on qubits, through second quantization and fermion-to-qubit encodings, and why the number of Hamiltonian terms matters so much. Chapter 3 gives the minimum of quantum computing a chemist needs: qubits, gates, circuits, measurement, and noise, with particular attention to what "exponential" does and does not mean. Chapter 4 develops the variational quantum eigensolver in full, from the variational principle through the main families of trial circuits and the classical optimizers that drive them, using molecular hydrogen as a worked example. Chapter 5 examines the three problems that stop VQE from scaling: the measurement cost, the effect of noise and the real price of error mitigation, and the barren plateau phenomenon that makes large circuits untrainable, together with the uncomfortable recent result linking trainability to classical simulability. Chapter 6 walks through the experimental record from 2014 to 2026, reading each landmark for what it did and did not show. Chapter 7 turns to the classical side: dequantization, tensor-network simulation of quantum experiments, and the evidence that classical heuristics scale better on chemical problems than was assumed. Chapter 8 describes the fault-tolerant route through quantum phase estimation, the resource estimates for flagship molecules and how quickly they have fallen, and the state preparation problem that no amount of hardware solves by itself. Chapter 9 turns all of this into practical judgment: how to read a claim, which problems are worth watching, and what a computational chemist should be doing now. A word on units and conventions. Energies are in hartree unless otherwise stated, and "chemical accuracy" means the conventional threshold of 1 kilocalorie per mole, about 1.6 millihartree. That threshold describes agreement with experiment for energy differences; many quantum computing papers use it instead to mean agreement with an exact calculation in a given basis, which is a much weaker claim. The distinction recurs throughout the book, because it is one of the easiest ways for a modest result to sound like a large one. The field moves quickly, and some of what follows will date. The reasoning should not. Whatever hardware exists when you read this, the questions that decide whether it helps a chemist are the same: how large is the problem, how is the state prepared, how many measurements does the answer need, how are errors handled, and what does the best classical method get for the same effort. Chapter 1: The Correlation Problem, Stated Honestly Any serious claim that a quantum computer will help chemistry has to begin with a precise account of what classical computers already do well, and where they genuinely struggle. Too much writing on quantum chemistry begins instead with the exponential size of the many-electron Hilbert space, as though the entire discipline of electronic structure theory were still waiting for someone to notice the problem. It is not. The exponential wall is real, but the field has spent seven decades finding ways around it, and those ways work for the great majority of molecules chemists care about. The question for quantum computing is not whether the exact problem is hard. It is whether there are molecules that matter, for which every practical classical approximation fails, and which a quantum algorithm could treat at a cost that is actually achievable. What correlation is, and why it resists The Hartree-Fock method treats each electron as moving in the average field of all the others. Its wavefunction is a single Slater determinant, an antisymmetrized product of one-electron orbitals. For a typical closed-shell organic molecule near its equilibrium geometry, Hartree-Fock recovers more than 99 percent of the total electronic energy. The remaining fraction, the correlation energy, is small in absolute terms and enormous in chemical terms. Reaction energies, barrier heights, binding energies and spectroscopic gaps are all differences between large numbers, and the correlation energy is often of the same order as the difference being computed. A method that ignores it cannot be trusted for chemistry. It is useful to divide correlation into two kinds, even though the division is not sharp. Dynamic correlation arises from the instantaneous tendency of electrons to avoid each other, which a mean field smooths away. It is spread thinly across very many small contributions from excited configurations, and it is present in every molecule. Static, or nondynamic, correlation arises when two or more electronic configurations are nearly degenerate, so that no single determinant is even qualitatively right. It appears when bonds are stretched or broken, in diradicals, in many excited states, and above all in molecules containing several open-shell transition metal ions whose d electrons couple in intricate ways. The two kinds call for different tools. Dynamic correlation is handled well by methods that start from one good reference determinant and add excitations systematically, such as perturbation theory and coupled cluster. Static correlation needs a multiconfigurational starting point: a wavefunction that treats a chosen set of orbitals, the active space, with a flexible expansion over many determinants. The hard problems of chemistry are the ones that need both at once: a large, strongly correlated active space, and dynamic correlation on top of it. The exact problem and its size In a finite basis set, the exact answer is full configuration interaction, a linear combination of every Slater determinant that can be formed by distributing the electrons among the orbitals. Its energy is the lowest eigenvalue of the Hamiltonian matrix in that determinant basis. Nothing is approximated except the basis itself. The number of determinants is what makes full configuration interaction impossible beyond small systems. For a singlet state with equal numbers of up-spin and down-spin electrons, the count is the square of the binomial coefficient choosing half the electrons from the number of spatial orbitals. Ten electrons in ten orbitals gives about 63,500 determinants, which a laptop handles instantly. Eighteen electrons in eighteen orbitals gives about 2.4 billion, which was the practical limit of conventional complete active space methods for many years. Twenty-six in twenty-six gives about 10 to the 14th. The active space used in an influential 2017 study of the nitrogenase cofactor, 54 electrons in 54 orbitals, corresponds to roughly 4 times 10 to the 30th determinants. Storing one number per determinant at that size would exceed all the digital storage on Earth by many orders of magnitude. This is the exponential wall, and it is the entire basis of the quantum computing argument. But the size of a space says nothing about how much of it the true wavefunction actually uses, and here chemistry has been fortunate. In most molecules the ground-state wavefunction is dominated by a small number of configurations, or has a structure, such as locality or low entanglement, that can be exploited. Nearly every successful classical method works by finding and exploiting such structure. What the classical toolbox actually achieves It is worth being specific about the main families, because each defines a different kind of competition for a quantum algorithm. Table 1 summarizes their formal cost and their characteristic strengths and weaknesses. Table 1. Main classical methods for correlated electronic structure. Method Formal cost Handles well Characteristic weakness Density functional theory Roughly cubic to quartic Large systems, geometries, trends Uncontrolled errors for strong correlation CCSD(T) and local variants Seventh power (local: near linear) Single-reference dynamic correlation Fails for bond breaking and multireference cases CASSCF with perturbation theory Exponential in active space Small strongly correlated active spaces Active space limited to about 20 orbitals exactly DMRG Polynomial in bond dimension Large active spaces, quasi-linear systems Cost rises with entanglement in 3D-like systems Selected CI (HCI and relatives) Grows with determinants kept Compact multireference wavefunctions Memory limits for very large active spaces AFQMC Roughly cubic to quartic per sample Dynamic plus some static correlation Phaseless bias depends on trial wavefunction Density functional theory deserves a word first because it is the method most chemists actually use. It is not systematically improvable, and its errors for strongly correlated systems are uncontrolled, but it is cheap enough to apply to thousands of atoms and good enough, with a well-chosen functional, for a remarkable range of questions. For many industrial applications the realistic competitor to a quantum computer is not an exact method at all but a density functional calculation that is fast, familiar, and usually adequate. Coupled cluster with single, double, and perturbative triple excitations, CCSD(T), is the reference method for molecules dominated by one configuration. Its canonical cost scales as the seventh power of system size, which sounds prohibitive, but local correlation methods such as domain-based local pair natural orbital coupled cluster have brought near-linear scaling to molecules with hundreds of atoms, with errors typically well under a kilocalorie per mole relative to canonical results. When the reference determinant is good, CCSD(T) is so reliable that it serves as the benchmark against which cheaper methods are judged. Its weakness is exactly the static correlation problem. Stretch the triple bond of dinitrogen and CCSD(T) first overestimates the energy, then turns over and produces an unphysical hump or collapses below the true curve, because the single reference has become qualitatively wrong. The same failure afflicts many transition metal systems, where near-degenerate d orbitals make any single determinant a poor starting point. For static correlation, the traditional answer is the complete active space self-consistent field method, CASSCF, followed by a perturbative treatment of dynamic correlation outside the active space, as in CASPT2 or NEVPT2. It works well when the essential physics fits in roughly eighteen to twenty orbitals. Parallel implementations have pushed exact active spaces into the low twenties, but the exponential growth makes each additional pair of orbitals far more expensive than the last. The density matrix renormalization group, DMRG, changed that picture in the 2000s. It represents the wavefunction as a matrix product state, a chain of small tensors whose size, the bond dimension, controls the accuracy. For systems whose entanglement is limited, DMRG converges to near-exact results with polynomial cost, and it has been applied to active spaces of fifty to one hundred orbitals. It is the method that made a detailed classical treatment of the iron-sulfur clusters in biology possible, and it is the most important single reason that the frontier of classically tractable strong correlation has moved as far as it has. Selected configuration interaction methods, such as heat-bath configuration interaction, identify the important determinants iteratively, keep only those, and correct for the rest with perturbation theory. For many molecules a few million well-chosen determinants out of 10 to the 20th or more are enough for near-exact energies. Auxiliary-field quantum Monte Carlo, AFQMC, takes a different route, sampling the ground state stochastically with a cost per sample that scales polynomially. It controls the sign problem with a constraint defined by a trial wavefunction, which introduces a bias that is small when the trial is good. On challenging transition metal benchmarks, AFQMC with a multideterminant trial has reached accuracy comparable to the best available alternatives. These methods are complementary, and practitioners combine them: a DMRG or selected CI active space, embedded in a larger coupled cluster or density functional calculation, with AFQMC or perturbation theory for the remaining correlation. The frontier is not a single method but a toolkit, and a quantum algorithm has to beat the best combination for the problem at hand, not the worst member of it. A case study in classical persistence The chromium dimer shows how the classical frontier actually moves. Two chromium atoms, each with six valence electrons in 3d and 4s orbitals, form a formal sextuple bond. The molecule looks trivially small, and for decades it was a notorious failure of quantum chemistry. Its potential energy curve has an unusual shoulder at intermediate bond lengths, and method after method produced curves that were qualitatively wrong: coupled cluster collapsed, multireference perturbation theory depended sensitively on the active space, and density functionals scattered widely. Chromium dimer became a standard test that the field used to show how hard strong correlation could be, and it was repeatedly cited as the kind of problem where only an exact solver would settle the question. It was settled classically. In 2022 a team using a combination of large-scale DMRG and selected configuration interaction, with careful extrapolation and perturbative corrections for dynamic correlation outside the active space, produced a potential energy curve in close agreement with the experimental shape and binding energy, and described the work as closing a chapter of quantum chemistry. The calculation was expensive and required expert judgment at every stage, but it did not require a new kind of computer. The lesson generalizes. Problems that are held up as intractable at one moment often become tractable through algorithmic improvement, and the rate of that improvement in classical electronic structure has been high. Anyone arguing that a specific molecule needs a quantum computer has to argue not only that current classical methods fail for it but that they will keep failing over the decade it would take to build the quantum machine. Where classical methods genuinely struggle Given all this, what is left? The honest list is shorter than the quantum computing literature often implies, but it is not empty. The first category is multinuclear transition metal clusters with many strongly coupled open-shell centers. The iron-molybdenum cofactor of nitrogenase, often called FeMoco, is the celebrated example. It contains seven iron atoms, one molybdenum, nine sulfurs and a central carbon, and it catalyzes the reduction of dinitrogen to ammonia at ambient conditions, a reaction that industrial chemistry performs at high temperature and pressure through the Haber-Bosch process. Understanding its mechanism requires resolving many close-lying spin states and oxidation states, and the relevant energy differences are small. Active spaces thought to capture the essential physics run from fifty to more than one hundred orbitals. DMRG calculations have been performed at this scale, but whether they are converged, and whether the chosen active space and model are chemically adequate, remains debated. Photosystem II's manganese-calcium cluster, cytochrome P450's iron-porphyrin chemistry, and various synthetic multimetallic catalysts belong to the same family. The second category is systems where static and dynamic correlation are both large and entangled with each other, so that splitting the problem into an active space plus a perturbative correction is unreliable. Certain bond-breaking reactions on metal surfaces, some lanthanide and actinide complexes, and strongly correlated extended materials fall here. For materials in particular, such as the doped copper oxide superconductors, the physics is often captured by lattice models like the Hubbard model, which have their own well-developed classical toolkit and their own open questions. The third category is not about ground states at all but about dynamics and excited states: real-time electron dynamics, nonadiabatic processes and spectroscopy. This booklet is about ground states, but it is worth noting that some of the strongest theoretical arguments for quantum advantage concern time evolution rather than ground-state energies, because there is no variational principle for dynamics and classical methods have fewer tricks to exploit. Even within these categories, it is important to be precise about what "struggle" means. Classical methods rarely fail to produce any answer. They produce answers whose error bars are uncertain, whose convergence is hard to demonstrate, or whose cost grows steeply enough that a converged answer is out of reach for the specific molecule and active space the chemist wants. A quantum advantage in this setting would not look like a quantum computer solving a problem that classical methods cannot touch. It would look like a quantum computer delivering a more certain answer, faster or more cheaply, on a problem where classical estimates disagree. What a useful answer requires The last piece of groundwork concerns what counts as a useful answer, because it shapes the precision a quantum computer must achieve. Chemistry almost never needs total energies. It needs energy differences: between reactants and products, between transition states and minima, between spin states. These differences are typically in the range of a few to a few tens of kilocalories per mole. To be useful, a calculation generally needs to get them right to about one kilocalorie per mole, 1.6 millihartree, the conventional threshold of chemical accuracy. Some questions, such as spin-state ordering in iron complexes, demand even finer resolution. This requirement has three consequences for quantum computing. First, a calculation that reproduces the exact energy in a minimal basis set is not chemically accurate in any meaningful sense, because the basis set error in a minimal basis is often tens or hundreds of millihartree. Agreement with full configuration interaction in the same small basis tests the quantum algorithm, not the chemistry. Second, the precision required for total energies is set by the energy differences of interest, and because errors do not always cancel between different geometries or spin states, each energy must usually be computed to better than the target difference. Third, and most importantly for later chapters, the cost of many quantum algorithms scales with the inverse of the precision, or its square. Tightening the target from ten millihartree to one millihartree can multiply the number of measurements a hundredfold. There is also the matter of the model. A quantum computer, like any exact solver, computes the exact answer within an active space and basis that a human or a classical procedure has chosen. If the active space omits orbitals that matter, or the embedding into the surrounding molecule and solvent is poor, an exact solution of the wrong model is still wrong. This is one reason that realistic proposals for quantum chemistry now describe hybrid workflows in which the quantum computer solves a carefully constructed active space problem while classical methods build the model, treat the environment, and add dynamic correlation from outside the active space. The shape of the opportunity Put together, these observations define a narrow but real target. A quantum computer helps a chemist when it can solve a strongly correlated active space that is too large or too entangled for DMRG, selected CI, or AFQMC to converge with confidence; when that active space, embedded properly, answers a chemical question that matters; and when the quantum calculation reaches the needed precision at a cost, in time and money, that competes with pushing the classical methods harder. Each condition narrows the field. Many strongly correlated active spaces that sound intimidating are in fact within reach of modern DMRG. Many chemically important questions can be answered by density functional theory with a sensible functional, or by coupled cluster, without ever needing a multireference solver. And the cost of a quantum calculation, as later chapters show, depends heavily on precision, measurement strategy, and error correction overhead. None of this means the opportunity is illusory. It means it lives where classical methods are weakest, and that its size can only be judged by specific comparison on specific molecules. The rest of the book follows that principle: every quantum method is measured against what the best classical method achieves for the same problem, and every demonstration is read for what it proves about that comparison. The next step is to see how the chemist's Hamiltonian becomes something a quantum computer can act on, because the choices made in that translation shape everything that follows. Chapter 2: From Orbitals to Qubits A quantum computer does not know what an electron is. It manipulates qubits, two-level systems with no charge, no spin in the chemical sense, and no antisymmetry. Before any algorithm can compute a molecular energy, the electronic Hamiltonian must be rewritten as an operator acting on qubits. This translation looks like bookkeeping, and much of it is, but the choices made here decide how many qubits a problem needs, how long the circuits are, and how many measurements the energy will cost. Several of the most important advances in quantum chemistry algorithms over the past decade have been improvements in this translation rather than in the algorithms that use it. Second quantization in brief Computational chemists usually meet the molecular Hamiltonian in first quantization, as a sum of kinetic energy operators, electron-nuclear attractions and electron-electron repulsions written in terms of electron coordinates. The antisymmetry of the wavefunction is imposed by writing it as a Slater determinant or a combination of them. Second quantization moves the antisymmetry out of the wavefunction and into the operators. One starts from a finite set of spin orbitals, typically the canonical Hartree-Fock orbitals or some localized variant. Each spin orbital is either occupied or empty, so any Slater determinant is fully described by a string of zeros and ones, its occupation number vector. For each spin orbital there is a creation operator, which places an electron in that orbital if it is empty, and an annihilation operator, which removes one if it is present. These operators anticommute: swapping the order of two creation operators on different orbitals changes the sign of the result. That single algebraic rule enforces the Pauli exclusion principle and the antisymmetry of every state built from them. In this language, the electronic Hamiltonian within the chosen orbital basis takes a compact form. It is a sum of one-electron terms, each a coefficient times a creation operator on one orbital and an annihilation operator on another, plus two-electron terms, each a coefficient times two creation and two annihilation operators. The one-electron coefficients are integrals of the kinetic and nuclear attraction operators over pairs of orbitals. The two-electron coefficients are the familiar electron repulsion integrals over four orbitals. Any standard quantum chemistry package, such as PySCF, Psi4 or Molpro, computes these integrals routinely, and they are the only molecular input a quantum algorithm needs. The number of two-electron integrals grows as the fourth power of the number of orbitals. With symmetry and the vanishing of integrals between distant orbitals the practical number is smaller, but the fourth-power growth is the natural scale of the problem. This matters because, as later chapters show, several quantum algorithms have costs that scale with the number of terms in the Hamiltonian or with the sum of the absolute values of their coefficients. The fermion-to-qubit problem Occupation number vectors look tailor-made for qubits. A qubit has two basis states, conventionally labeled zero and one, and a register of qubits has basis states labeled by bit strings. Map each spin orbital to one qubit, and let the qubit's state record whether the orbital is occupied. A Slater determinant becomes a computational basis state of the qubit register. The Hartree-Fock determinant of a molecule with a given number of electrons is simply the bit string with ones in the lowest occupied orbitals and zeros elsewhere. The difficulty is the operators. Qubits do not anticommute. Operators on different qubits commute with each other, whereas fermionic operators on different orbitals anticommute. A faithful mapping must reproduce the fermionic sign structure using only qubit operations. Three standard encodings do this in different ways, and their trade-offs are summarized in Table 2. Table 2. Common fermion-to-qubit encodings. Encoding Stores in each qubit Operator weight Main advantage Main drawback Jordan-Wigner Occupation of one orbital Up to number of qubits Simple and transparent Long strings of Z operators Parity Parity of all orbitals up to it Up to number of qubits Two qubits removable by symmetry Long strings for creation operators Bravyi-Kitaev Partial sums of occupations Logarithmic in qubits Shorter operators at scale Less intuitive, gains modest for small molecules The Jordan-Wigner transformation, dating to 1928 and originally developed for spin chains, is the most transparent. Each qubit stores the occupation of one spin orbital. A creation or annihilation operator on orbital number k becomes a raising or lowering operation on qubit k, multiplied by a string of Pauli Z operators on all qubits with lower index. The Z string counts, with a sign, how many occupied orbitals precede orbital k, which is exactly what the anticommutation rule requires. The price is that an operator involving a high-index orbital acts nontrivially on many qubits, and on hardware where qubits interact only with their neighbors, long strings translate into long circuits. The parity encoding makes the opposite choice. Each qubit stores the parity, even or odd, of the total occupation of all orbitals up to and including its own. Parity information that the Jordan-Wigner scheme must compute with a long Z string is now available locally, but updating an occupation requires flipping every later qubit, so the long strings reappear elsewhere. Its practical virtue is symmetry reduction: the last qubit stores the total electron number parity, and with a suitable ordering another qubit stores the parity of one spin species. Both are fixed for a given molecule, so two qubits can be removed from the problem with no loss. The Bravyi-Kitaev encoding, introduced in 2002 and brought into chemistry around 2012, interpolates between the two, storing partial sums of occupations arranged on a binary tree. Both occupation and parity information then require acting on only a logarithmic number of qubits. For large systems the reduction in operator weight is significant in principle. For the small molecules that have actually been run on hardware, the difference is usually modest, and the choice of encoding is often driven by compatibility with the circuit design rather than asymptotic scaling. A worked example: molecular hydrogen The standard first example is the hydrogen molecule in the minimal STO-3G basis. Each hydrogen contributes one 1s function, giving two spatial molecular orbitals, the bonding and antibonding combinations, and four spin orbitals. Under the Jordan-Wigner transformation, four spin orbitals become four qubits, and the Hartree-Fock state, with both electrons in the bonding orbital, is the bit string with ones on the two bonding spin orbitals. The Hamiltonian becomes a weighted sum of fifteen Pauli strings, counting the identity, each a product of single-qubit Pauli operators X, Y, Z or the identity on each of the four qubits. The coefficients depend on the bond length through the integrals. The energy at any bond length is the lowest eigenvalue of this 16 by 16 matrix, restricted to states with two electrons and zero spin projection. Using symmetries, including electron number, spin, and the spatial symmetry that forbids single excitations between orbitals of different parity, the problem can be reduced to two qubits, and even to one: the only determinants that mix in the singlet ground state are the doubly occupied bonding configuration and the doubly occupied antibonding configuration, so a single rotation angle describes the exact answer. That reduction is illuminating in two ways. It shows how much of a naive qubit count can be removed by symmetry, which matters for fitting problems onto small devices. And it shows why hydrogen in a minimal basis is a test of plumbing, not chemistry. It has one parameter, a two-by-two eigenvalue problem, and an answer that is known exactly in closed form. Every early hardware demonstration was, at heart, a demonstration that the plumbing worked. Qubit counts in realistic problems With one qubit per spin orbital, qubit count is twice the number of spatial orbitals in the active space. A strongly correlated active space of 50 spatial orbitals needs 100 qubits before symmetry reduction. The full cofactor models used for nitrogenase, with 54 or 76 active orbitals, need 108 or 152 system qubits. These numbers are sometimes quoted as though they were the size of the quantum computer required. They are not: they count logical system qubits only, and any practical algorithm needs additional qubits for ancillas, for data loading, and above all for error correction, which can multiply physical qubit counts by factors of hundreds or thousands. The basis set question is where many demonstrations quietly part company with chemistry. A calculation in a minimal basis captures qualitative bonding but not quantitative energetics. Chemically useful results typically need at least a correlation-consistent double-zeta basis and often triple-zeta, with the number of orbitals growing rapidly. Hydrogen in cc-pVDZ has ten spatial orbitals, twenty spin orbitals, and already needs twenty qubits. A modest organic molecule in a triple-zeta basis has hundreds of orbitals. No quantum algorithm treats all of them directly. Instead, as with classical multireference methods, the quantum computer treats an active space, and the orbitals outside it are handled by classical methods: frozen, treated perturbatively, or folded in through embedding schemes or transcorrelated Hamiltonians that absorb some dynamic correlation into a modified operator. This makes the quantum computer an active-space solver, a drop-in replacement for the configuration interaction or DMRG step in a CASSCF-style calculation. It is a useful way to think about the whole enterprise, because it immediately identifies the classical competitor: the best available active-space solver for the same active space. It also identifies a hidden cost. Orbital optimization in CASSCF requires not just the energy but the one- and two-particle reduced density matrices of the active-space wavefunction, and extracting those from a quantum computer requires many more measurements than the energy alone. Symmetry, tapering and conserved quantities The hydrogen example hinted at a general tool. Molecular Hamiltonians commute with several symmetry operators: total electron number, the number of electrons of each spin, total spin, and the point group of the nuclear framework. Each conserved quantity divides the full space of qubit states into sectors, only one of which contains the state of interest. A calculation that wanders into the wrong sector, for instance one that ends up with three electrons in a two-electron problem because of hardware noise, produces a meaningless energy. Some of these symmetries appear, after the qubit mapping, as products of Pauli Z operators that commute with every term of the Hamiltonian. A procedure introduced in 2017, often called qubit tapering, finds such symmetries automatically and applies a transformation that turns each one into a single-qubit operator. That qubit's value is then fixed by the chosen symmetry sector, and it can be deleted. For small molecules with high point-group symmetry, tapering removes several qubits at no cost in accuracy, which is why published experiments sometimes treat a molecule on fewer qubits than its spin orbital count would suggest. Symmetries that cannot be tapered away can still be protected in other ways. Circuits can be built entirely from gates that conserve particle number and spin projection, so that the state never leaves the right sector in the absence of noise. When noise does push it out, measuring the conserved quantity and discarding runs that violate it, a technique called symmetry verification or postselection, removes a portion of the errors at the cost of throwing away data. That trade, fewer errors for more measurements, is typical of what the next chapters call error mitigation. Choosing and ordering orbitals The orbitals used in the mapping are not dictated by the molecule. Canonical Hartree-Fock orbitals are delocalized over the whole molecule. Localized orbitals, obtained by rotations such as the Pipek-Mezey or Foster-Boys procedures, concentrate each orbital on a few atoms. Natural orbitals, the eigenfunctions of a correlated one-particle density matrix, order themselves by occupation and give the most compact configuration interaction expansions. Every choice describes the same Hamiltonian, but they distribute its complexity differently. For quantum algorithms this matters in two ways. Under the Jordan-Wigner transformation, the length of each Pauli Z string depends on the order in which orbitals are assigned to qubits, so an ordering that places strongly interacting orbitals next to each other shortens circuits, particularly on hardware where only neighboring qubits interact. And the choice of orbitals determines how much of the correlation must be described by the quantum state at all: a good orbital basis puts the wavefunction closer to a simple starting point and reduces the work the quantum circuit has to do. DMRG practitioners know both effects well, because the same orbital ordering and localization choices govern the bond dimension a matrix product state needs. It is another instance of classical expertise transferring directly to the quantum setting. Factorizing the Hamiltonian Because so many costs scale with the number and size of Hamiltonian terms, a large body of work has focused on rewriting the two-electron part more efficiently. The idea is borrowed from classical quantum chemistry, where the Cholesky decomposition and density fitting of the electron repulsion integrals have been standard for decades. The four-index tensor of two-electron integrals can be viewed as a matrix between pairs of orbital indices, and that matrix is typically of low rank: its eigenvalues fall off quickly, so a truncated decomposition reproduces it well. A single factorization expresses the two-electron operator as a sum of squares of one-body operators. A double factorization goes further, diagonalizing each of those one-body operators so that each term becomes, in a rotated orbital basis, a simple function of occupation numbers. On a quantum computer, orbital rotations can be implemented efficiently with circuits of Givens rotations, and occupation-number functions are cheap to evaluate or measure. The result is a Hamiltonian expressed as a modest number of groups, each measurable in a single rotated basis or simulable with a single structured circuit. Tensor hypercontraction, also borrowed from classical electronic structure, pushes the compression further by approximating the integral tensor with a product of smaller factors. In 2021 a Google-led team used it to produce what were then the lowest published resource estimates for fault-tolerant simulations of the nitrogenase cofactor, a result discussed in Chapter 8. These methods matter for near-term algorithms as well, because they reduce the number of distinct measurement settings a variational calculation needs. First quantization as an alternative Everything above assumes a second-quantized description with a fixed orbital basis, which is natural for chemists and dominant in near-term work. There is an alternative. In first quantization, each electron is assigned a register of qubits that stores its position on a grid, or its index in a large basis set such as plane waves. The number of qubits then scales with the number of electrons times the logarithm of the basis size, rather than with the basis size itself. Antisymmetry must be imposed on the initial state, but the Hamiltonian operations preserve it thereafter. For very large basis sets, first quantization can be dramatically more economical in qubits, and plane waves avoid the basis set incompleteness of Gaussian orbitals in a systematic way. Resource estimates for fault-tolerant simulation in first quantization have been published for materials and for molecules in the plane-wave basis, and they suggest this route may be preferred for some problems once large error-corrected machines exist. For near-term devices, the additional arithmetic required makes it impractical, and nearly all experimental work has used second quantization. What the mapping teaches Three lessons from this chapter recur throughout the book. First, the qubit count of a chemical problem is set by the active space, not the molecule, and headline qubit counts are always a statement about a chosen model. Second, the structure of the Hamiltonian, meaning its number of terms, their magnitudes, and how they can be grouped or factorized, directly determines the cost of both near-term and fault-tolerant algorithms; improvements in that structure have repeatedly produced larger practical gains than improvements in hardware. Third, the translation is where chemistry and quantum information meet most concretely, and it rewards chemists who understand both. The same integrals, the same orbital choices and the same symmetries that make classical methods efficient make quantum methods efficient too. With the Hamiltonian expressed as an operator on qubits, the next question is what a quantum computer can do with it. That requires a short but careful account of the machine itself. Chapter 3: The Machine, for Chemists A chemist does not need to understand superconducting circuit design or ion trap optics to evaluate quantum chemistry algorithms. But a few features of how quantum computers work are decisive for everything that follows, and they are often explained either too loosely to be useful or too formally to be read. This chapter gives the minimum, framed wherever possible in terms a computational chemist already uses. States, amplitudes and the catch A single qubit is a two-level quantum system. Its state is a superposition of the two basis states, zero and one, described by two complex amplitudes whose squared magnitudes sum to one. So far this is familiar: it is the same mathematics as a spin one-half particle, or a two-state configuration interaction problem. A register of n qubits has two to the n basis states, one for every bit string, and its general state is a superposition over all of them, described by two to the n complex amplitudes. Under the Jordan-Wigner mapping of the last chapter, those bit strings are Slater determinants and those amplitudes are configuration interaction coefficients. A register of 100 qubits can therefore hold a wavefunction over 2 to the 100th determinants, about 10 to the 30th, which is precisely the full configuration interaction space of a large strongly correlated active space. This is the central attraction: the quantum computer does not store the wavefunction as a list of numbers; it physically is the wavefunction. The catch is equally central. The amplitudes cannot be read out. When a qubit register is measured in the computational basis, the result is a single bit string, drawn at random with probability equal to the squared magnitude of its amplitude, and the superposition is destroyed. To learn anything about the state, one must prepare it again and measure again, many times, and infer properties from the statistics. A quantum computer that holds a perfect full configuration interaction wavefunction for FeMoco cannot simply print the coefficients. It can only be sampled. For chemistry this means the output of any quantum algorithm must be a quantity that can be estimated efficiently from samples: an energy, a dipole moment, a reduced density matrix element, or a phase. Many useful quantities qualify. But the number of samples needed to reach a given precision becomes a first-order cost, and it is where near-term algorithms most often run into trouble. Gates and circuits Quantum computations are built from gates, elementary unitary operations on one or two qubits, arranged in sequence into circuits. Single-qubit gates rotate a qubit's state; two-qubit gates, such as the controlled-NOT, entangle pairs. A small set of such gates is universal, meaning that any unitary operation on the register can be approximated to arbitrary accuracy by a sequence of them. Universality says nothing about efficiency. Most unitaries on n qubits require a number of gates that grows exponentially with n. The useful ones, those that implement a quantum algorithm with a polynomial number of gates, have special structure. For chemistry, the most important structured operations are orbital rotations, which transform one set of orbitals into another and can be implemented with circuits of Givens rotations whose depth grows linearly with the number of orbitals; time evolution under the molecular Hamiltonian, which can be approximated by products of exponentials of individual terms; and exponentials of excitation operators, the quantum counterparts of the excitation operators in coupled cluster theory. Two properties of a circuit dominate its practical cost. The gate count, particularly the number of two-qubit gates, determines how much error accumulates, because two-qubit gates are typically an order of magnitude noisier than single-qubit ones. The depth, the number of sequential layers of gates, determines how long the qubits must hold their state, and so how much they decohere. Both depend on hardware connectivity. If only neighboring qubits can interact directly, a two-qubit gate between distant qubits must be built from a chain of swaps, inflating both count and depth. A small circuit, read as chemistry It helps to trace one tiny circuit in chemical terms. Take the four-qubit hydrogen problem from the last chapter, with qubits zero and one representing the up and down spin orbitals of the bonding orbital, and qubits two and three the antibonding ones. All qubits start in the zero state, the empty vacuum. Applying a bit-flip gate, the Pauli X, to qubits zero and one produces the bit string with ones in the bonding spin orbitals: the Hartree-Fock determinant, prepared exactly with two single-qubit gates. Now apply a structured sequence of two-qubit gates that implements a rotation by some angle between the doubly occupied bonding configuration and the doubly occupied antibonding configuration. The register now holds a superposition of two determinants, with coefficients equal to the cosine and sine of that angle. This is a two-configuration wavefunction of exactly the kind a chemist would write down to describe a stretched hydrogen bond, and for hydrogen in this basis it contains the exact ground state for a suitable angle. Measuring the register in the computational basis returns either the Hartree-Fock bit string or the doubly excited one, with probabilities given by the squared coefficients. Averaging many such measurements gives the expectation value of any Hamiltonian term built only from Z operators, which in the Jordan-Wigner picture correspond to orbital occupations and their products. The Hamiltonian also contains terms with X and Y operators, which correspond to the exchange-like processes that move electrons between orbitals. Those cannot be read off in the computational basis. Before measuring them, one applies single-qubit rotations that change the measurement basis, so that each such term becomes diagonal. Different groups of terms need different rotations, and so different runs of the circuit. That small example contains the whole structure of the variational approach in miniature: prepare a reference, apply a parameterized entangling circuit, measure in several bases, and assemble an energy from the averages. It also contains its principal cost, which is that every expectation value comes from repeated preparation and measurement. Today's hardware Three physical platforms dominate current machines, and their differences matter for chemistry. Superconducting qubits are small electrical circuits cooled to about ten millikelvin. Gates take tens of nanoseconds, so circuits run quickly, but each qubit connects only to a few neighbors on a fixed chip layout. Google's Willow processor, announced in December 2024, has 105 qubits. IBM's Heron processors, with 133 and later 156 qubits, have been the workhorses of its cloud service, and in November 2025 IBM introduced Nighthawk, a 120-qubit chip with a square lattice in which each qubit couples to four neighbors, designed to support circuits with around 5,000 two-qubit gates. Trapped-ion qubits are individual charged atoms held in electromagnetic traps and manipulated with lasers. Gates are slower, typically tens to hundreds of microseconds, but fidelities are the highest of any platform and ions can be shuttled so that any pair can interact. Quantinuum's Helios system, launched commercially in November 2025 and described in Nature in 2026, holds 98 barium ions with all-to-all connectivity, and the published characterization reports average two-qubit gate infidelities of about 8 in 10,000, a fidelity near 99.92 percent. Neutral atom machines hold uncharged atoms in arrays of optical tweezers and entangle them by exciting them to high-energy Rydberg states. They scale in qubit number more easily than the other platforms, with arrays of thousands of atoms demonstrated, and atoms can be rearranged during a computation. Their gate fidelities have improved rapidly, and some of the most ambitious logical qubit experiments of the past few years have been done on this platform. A fair summary for chemists is this: the best physical two-qubit gates now fail somewhere between one time in a thousand and one time in a few hundred, depending on platform and operating conditions, and the largest machines have between about one hundred and a few thousand physical qubits. Neither number, by itself, says whether a chemical calculation is feasible. What matters is how many gates the calculation needs relative to the inverse error rate, and how the errors are handled. Why classical simulators run out The exponential size of the state space explains why quantum circuits cannot simply be simulated classically at scale, and it also explains where the line sits. A statevector simulator stores all two to the n complex amplitudes, each taking sixteen bytes in double precision. Thirty qubits need about 17 gigabytes, within reach of a workstation. Forty qubits need about 17 terabytes, which requires a large cluster. Fifty qubits would need about 18 petabytes, more memory than any existing supercomputer possesses. Brute-force simulation therefore stops somewhere in the mid-forties of qubits. That boundary is often quoted as the point beyond which a quantum computer does something classically impossible. It is not. Tensor-network simulators do not store the full state; they store a compressed representation whose size depends on how much entanglement the circuit creates, and for circuits with limited entanglement they can simulate hundreds of qubits. Other methods propagate observables rather than states and exploit the damping effect of noise. The relevant boundary is not the qubit count at which brute force fails but the circuit complexity at which every clever classical method fails, and that boundary is much harder to locate. It is the same distinction as the one between exact diagonalization and DMRG in electronic structure, and it will recur throughout the book. Noise and the error budget Every gate, every idle moment and every measurement introduces a small probability of error. Errors accumulate. A useful rule of thumb is that a circuit can reliably execute roughly as many two-qubit gates as the inverse of the two-qubit error rate before the output is dominated by noise. With an error rate of one in a thousand, that is on the order of a thousand gates, spread across all the qubits. More gates than that and the probability that the whole circuit ran without any error becomes small, and the computed quantity drifts toward the value it would have in a random state. For chemistry this is a severe constraint. A chemically meaningful ansatz for an active space of even twenty orbitals, built from coupled-cluster-style excitation operators and compiled for hardware with limited connectivity, can easily need tens of thousands of two-qubit gates or more. Time evolution for phase estimation on problems of interest needs many orders of magnitude more. John Preskill named the current period the era of noisy intermediate-scale quantum devices, NISQ, in 2018, precisely to mark the gap between what these devices can hold and what they can reliably compute. There are two responses to noise. Error mitigation accepts that the hardware is noisy and uses classical post-processing of carefully designed extra experiments to estimate what the noiseless result would have been. It needs no extra qubits, but its cost in samples grows, in general exponentially, with the size of the circuit. Error correction encodes each logical qubit in many physical qubits, detects errors as they occur by measuring collective properties of the physical qubits, and corrects them before they spread. It needs large overheads in qubits and time, but it lets computations run arbitrarily long, provided the physical error rate is below a threshold. Chapter 5 examines mitigation and Chapter 8 examines correction; the distinction between them is the most important dividing line in the field. What "exponential" does and does not mean Popular accounts often say that a quantum computer with n qubits performs two to the n calculations at once. That description is misleading, and for chemists it leads to a specific error: the belief that because a quantum register can hold a full configuration interaction wavefunction, a quantum computer can find the ground state of any molecule efficiently. It cannot. The relevant result comes from computational complexity theory. Finding the ground-state energy of a general local Hamiltonian, even to modest precision, is QMA-complete, the quantum analogue of NP-complete. It is believed that no efficient algorithm, classical or quantum, solves it in the worst case. A 2009 result showed that the electronic structure problem in its general form, electrons in an arbitrary external potential, inherits this hardness. So a quantum computer does not make ground-state chemistry easy in general, any more than a classical computer makes the traveling salesman problem easy. What a quantum computer can do efficiently is somewhat different. Given a Hamiltonian and an initial state that has a reasonable overlap with the true ground state, quantum phase estimation projects onto the ground state with probability equal to the squared overlap and returns its energy to a precision that improves linearly with the length of the computation. The hard part, finding the ground state, has been converted into a different hard part, preparing a state that already resembles it. If a simple state such as Hartree-Fock has large overlap with the true ground state, the problem is easy. If the overlap is exponentially small, the procedure takes exponentially many repetitions. This reframing is essential for everything that follows. The case for quantum advantage in ground-state chemistry rests on the existence of molecules in a middle ground: hard enough that classical methods cannot solve them with confidence, but structured enough that a good starting state can be prepared on a quantum computer. Whether such molecules exist, and how many, is not a question complexity theory answers. It is an empirical question about chemistry, and it is at the center of Chapter 7. Sampling and precision Because every useful output is estimated from repeated measurement, statistical precision follows the familiar rule of Monte Carlo methods. The standard error of an average over N independent samples falls as one over the square root of N. Halving the error requires four times as many samples; improving it tenfold requires a hundred times as many. For an energy computed by measuring each Hamiltonian term separately, the number of samples required to reach a target precision depends on the sum of the variances of the terms, which in turn depends on the magnitudes of the Hamiltonian coefficients. Molecular Hamiltonians have many terms, and some of their coefficients are large. The result, examined quantitatively in Chapter 5, is that estimating a molecular energy to one millihartree by direct sampling can require an enormous number of circuit repetitions, and that this cost, not the qubit count, is often the binding constraint on variational algorithms. Phase estimation escapes this particular scaling. Its precision improves as one over the total evolution time rather than one over the square root of the number of samples, a quadratic improvement known as the Heisenberg limit. That improvement is one of the main reasons the fault-tolerant route looks more promising than the variational one for problems that need chemical accuracy. It comes at the cost of circuits far deeper than any noisy device can run. Logical and physical qubits A final distinction is between physical and logical qubits. A physical qubit is a single ion, atom or superconducting circuit. A logical qubit is an error-corrected qubit encoded in many physical ones. The most studied code, the surface code, uses a two-dimensional patch of physical qubits whose size grows with the code distance; the logical error rate falls exponentially with distance provided the physical error rate is below the threshold. In December 2024 Google reported in Nature a surface code memory on its Willow processor in which the logical error rate fell by roughly a factor of two each time the code distance increased by two, from distance three to distance five to distance seven. This was the first convincing demonstration of below-threshold scaling in a superconducting surface code, and it is a genuine milestone. It was a memory experiment, preserving one logical qubit, not a computation, and the logical error rate at distance seven, around one in a thousand per cycle, is still far above what long chemistry algorithms need. Trapped-ion and neutral atom groups have demonstrated small numbers of logical qubits performing logical operations, and the first chemistry calculations on logical qubits, described in Chapter 6, have been run. The practical meaning for chemists is that resource estimates for fault-tolerant chemistry are quoted in logical qubits and logical operations, and translate into physical qubits through code overheads that depend on physical error rates, code choice and architecture. An algorithm needing a few thousand logical qubits may need a few million physical qubits under surface code assumptions, or fewer with newer codes and hardware. Those numbers, and how quickly they have been falling, are the subject of Chapter 8. First, though, comes the algorithm that was supposed to make all of this unnecessary for the near term. Hashtags: #QuantumComputingForComputationalChemists #QuantumChemistry #ComputationalChemistry #ElectronicStructure #ElectronCorrelation #MolecularGroundStates #VariationalQuantumEigensolver #VQE #QuantumPhaseEstimation #StrongCorrelation #ActiveSpaceMethods #FermionToQubitMapping #JordanWignerTransformation #BravyiKitaevTransformation #QubitTapering #MolecularHamiltonians #QuantumCircuits #ErrorMitigation #QuantumErrorCorrection #FaultTolerantQuantumComputing #StatePreparation #MeasurementCost #ChemicalAccuracy #QuantumAdvantage #FutureOfQuantumChemistry Pasted markdown

  • Qualitative Inquiry in the Digital Space (Ethnography, Discourse, and Network Mining)

    Download the Book (PDF): Introduction A researcher sits down to study how people who have been through a particular illness talk to one another. Twenty years ago this would have meant finding a support group, negotiating access with the facilitator, sitting in a circle for eight months, and writing fieldnotes on the drive home. Today it means something stranger. The group exists, but it exists as a forum with fourteen thousand members, a private messaging channel that splintered off from the forum after an argument in 2019, three overlapping hashtag publics on different platforms, a video channel where one member posts monologues that the others dissect, and a set of direct-message conversations the researcher will never see and should probably never want to. The methods literature offers two tempting responses to this situation, and both are wrong. The first is to treat the online setting as simply a new location for old methods. Ethnography moves to the forum; the interview happens over video; the fieldnotes get typed rather than scrawled. On this view nothing fundamental has changed, and the researcher's job is to apply Malinowski's injunction to grasp the native's point of view in a place that happens to have a URL. This response has the virtue of taking interpretation seriously. Its defect is that it underestimates how much the setting does to the data. A forum post is not an utterance overheard; it is a durable, searchable, platform-formatted, algorithmically ranked artefact composed for an audience the writer could only guess at, and possibly rewritten three times before posting. Reading it as though it were speech in a circle of chairs misses most of what is going on. The second response is to treat the online setting as a data source and reach for scale. Here the fourteen thousand members become four million posts, the posts become a matrix, the matrix gets a sentiment score and a topic model and a network layout, and the resulting patterns are reported as findings about how people who have been through this illness talk. This response has the virtue of taking the size of the archive seriously. Its defect is that the numbers it produces are often measurements of the platform rather than of the people. A spike in negative sentiment may record a change in who was posting, or a change in the ranking algorithm, or the arrival of a brigading campaign, or nothing more than the fact that a lexicon built on product reviews scores the word "positive" as good and so reads "my biopsy came back positive" as joy. This booklet argues for a third position, and it is not a compromise between the other two. The claim is this: qualitative inquiry online is distinguished by the work it does to keep interpretation attached to context, and that work has to be designed into the ethics, the access route, and the analytic method simultaneously, because each of those three decisions destroys or preserves a different part of the context the interpretation depends on. That is a single, demanding idea, and everything that follows is an unpacking of it. Consider what it rules out. It rules out doing the ethics at the end, as a compliance exercise, because the consent route you choose determines whether you can quote, which determines whether your reader can evaluate your reading. It rules out treating data acquisition as a technical preliminary, because whether you obtain material through an official research interface, a data donation, a public archive, or your own careful reading determines what metadata survives, what sampling frame exists, and what you are permitted to say about the population. It rules out bolting computational text analysis onto a qualitative study as a garnish, because a classifier trained on decontextualised strings and a close reading of a thread are not two views of the same object; they are answers to different questions, and the researcher has to know which question is being asked. Context is the word doing the heavy lifting here, so it is worth being precise about it. Helen Nissenbaum's account of privacy as contextual integrity gives us the sharpest available version: information flows carry norms attached to the context in which they occur, and a breach happens not when private information becomes public but when information moves from one context into another under different norms. A post in a recovery forum is not secret; it is contextually situated. Moving it into a journal article, or into a training corpus, or into a spreadsheet of sentiment scores, is a flow across contexts, and each such flow has to be justified rather than assumed. But context is also an epistemic matter, not only an ethical one. The meaning of a message depends on the thread it answers, the community's history of similar exchanges, the platform's affordances for sarcasm and quotation, the poster's standing in the group, and what everyone knows happened last week. Strip that away and you have a string of characters that can be counted but not understood. The characteristic error of weak digital qualitative work is not that it is too small or too impressionistic; it is that it takes decontextualised material and interprets it as though the context were still attached. What this booklet covers, and what it does not The chapters move roughly in the order a study moves. The first chapter takes on the question of what the field is when it has no location, and how researchers from Christine Hine onwards have answered it. This matters practically: your definition of the site is your sampling frame in disguise, and most weak designs are weak because the site was never defined at all. The second chapter puts ethics where it belongs, which is before method rather than after it. The public–private binary is useless online; contextual integrity, risk-based reasoning, and honest thinking about what a quotation can do to a person are what replace it. Chapters three and four cover the two classical field methods in their online forms: participant observation, where the difficulty is presence rather than access, and asynchronous interviewing, which is not a degraded version of a face-to-face interview but a genuinely different instrument with its own strengths. The fifth chapter addresses how a corpus is lawfully assembled. This is the least glamorous chapter and possibly the most useful one. It treats the API, the regulated research-access channel, the data donation, and the public archive as distinct instruments with distinct properties, and it is unapologetically clear that evading a platform's technical controls is not a method. Chapters six, seven, and eight cover analysis: close discourse analysis of online talk, computational text methods including sentiment analysis, and the mining of relational structure. The argument across those three chapters is that computational methods are best understood as ways of finding places to read closely, and that their outputs are hypotheses about a corpus rather than findings about people. The ninth chapter is about writing, which online research makes unusually hard because verbatim quotation is searchable and therefore identifying. Solutions exist; they involve trade-offs the reader of your work deserves to be told about. Several things are deliberately out of scope. This is not a book about how to circumvent platform controls, defeat rate limits, or re-identify individuals from de-identified data; where those topics arise they are treated as things researchers must not do and as risks that responsible design mitigates. It is not a statistics text, and it does not teach any particular software. It is not a survey of every platform, because platforms change faster than books are printed and a method tied to one interface has a shelf life measured in months. And it does not pretend that the four-year-old tooling landscape is stable: the closure of CrowdTangle in August 2024, the collapse of free academic access to what was then Twitter's API in 2023, and the arrival of a regulated European research-access regime under the Digital Services Act have all rearranged the practical terrain in ways this booklet describes but cannot freeze. Who this is for The reader I have in mind has done, or is about to do, empirical social research and has a reasonable grasp of what interpretation involves. They may be a doctoral student working out whether their project is feasible, a researcher in a policy or public-health setting who needs to understand online conversation without a computational team, or a computational social scientist who has begun to suspect that their models are answering a question nobody asked. What the booklet tries to give all three is judgement rather than procedure. There are procedures here — how to build a fieldnote practice when there is no drive home, how to structure an asynchronous interview so it does not die after the second exchange, how to validate a sentiment classifier against a hand-coded sample, how to decide whether a quotation is safe to print — but procedures without judgement produce studies that are correct in every step and worthless overall. The underlying stance is not sceptical about computation. Large-scale text and network methods have genuinely extended what qualitative researchers can see; a topic model that surfaces a register of talk you had not noticed in eleven thousand posts has done you a service no amount of close reading would have. The stance is sceptical about unvalidated computation, and about the particular illusion that the digital record is a transcript of social life rather than a trace of interaction with a commercial system that was designed for other purposes. Two commitments follow, and they run through every chapter. The first is that no quantitative output about text or ties is a finding until someone has read enough of the underlying material to say what it means. The second is that the ethical obligations of this work do not weaken as the data gets bigger. A research subject does not stop being a person because there are four million of them. That is the argument. The rest is the working out. Chapter 1: The Field Without a Location Bronisław Malinowski's famous instruction to the ethnographer was to pitch a tent in the village, learn the language, and stay long enough that the strange becomes ordinary. The instruction assumed something so obvious that he never had to say it: that there was a village, that it had edges, and that being inside those edges meant something. Almost every difficulty in online qualitative research traces back to the fact that this assumption no longer holds, and to the various unsatisfying ways researchers have tried to patch it. The patches matter because the definition of the field is not a philosophical preliminary. It is the sampling frame, the ethical perimeter, and the claim of the eventual write-up, all bundled together and usually left implicit. A study that never decided what its field was cannot say what its findings are about. When a reviewer asks "how do you know this is characteristic of the community rather than of the loudest twelve accounts?", they are asking a question about site definition that should have been answered before any data was collected. Four ways of drawing the boundary The literature has settled into roughly four framings, and they are not interchangeable. Each defines the field by a different principle, and each buys a different kind of validity at a different cost. The bounded-platform framing treats a particular forum, subreddit, server, or group as the site, in more or less the way the village was a site. This is what most early internet ethnography did, and it remains the right choice when the community really does have a membership, a shared history, and a moderating structure that constitutes it as a group. Tom Boellstorff's work in Coming of Age in Second Life (2008) is the most thoroughgoing defence of the position. He argued, against the reflex that virtual worlds are somehow less real, that a world with its own economy, norms, and social hierarchies can be studied in its own terms, without constant reference to the offline lives of the people playing. That is a methodological argument about boundaries: the field is where the action is, and the action can be inside the world. The networked-practice framing, associated above all with Christine Hine, denies that the site is a place at all. In Virtual Ethnography (2000) and more fully in Ethnography for the Internet (2015), Hine argued that the internet is at once a culture in its own right and a cultural artefact embedded in other settings, and that the ethnographer follows connections rather than staying put. The field becomes a trajectory: this forum, plus the Discord it spills into, plus the conference where members meet, plus the domestic contexts where people read it on their phones at midnight. George Marcus's older idea of multi-sited ethnography, of "following the thing" or "following the metaphor," is the obvious ancestor. The event or phenomenon framing defines the field by something happening rather than somewhere existing: a controversy, a hashtag surge, a moderation crisis, a product launch. Here the boundary is temporal and thematic. Much of what is published as social media research is implicitly of this kind, and it is a legitimate design provided the researcher acknowledges that they are studying an event's public trace and not a community. The person-centred framing puts a set of participants at the centre and treats the field as whatever media practices those people engage in. This is the design behind the UCL Why We Post project, published as Daniel Miller and colleagues' How the World Changed Social Media (2016), in which nine ethnographers spent fifteen months each in different locations and studied social media as one strand of local life rather than as a separate world. It is also the logic of most work that pairs interviews with participants' own accounts of their platform use. The finding such a design supports is about people; it cannot easily support claims about a platform in general. The choice among these is a choice about what you will be able to claim, as Table 1 sets out. Table 1. Four ways of defining the online field and what each supports. Framing Boundary principle Claims it supports Characteristic failure Bounded platform Membership in one forum or world Norms, roles, and internal culture of that group Treating one community's culture as "how people online behave" Networked practice Connections followed across settings How a practice travels and changes across contexts Drift; a field so diffuse that nothing is observed in depth Event or phenomenon A controversy or period of activity The shape and repertoire of a public episode Mistaking the trace of an event for a community Person-centred A set of participants and their media use How specific people integrate platforms into life Overreach from a small purposive sample to a platform None of these is more rigorous than the others. What is not rigorous is choosing none and letting the field be defined by whatever the search interface happened to return. The three things that make a digital site different Even once a boundary is drawn, the site has properties a physical one does not, and three of them bear directly on method. The first is persistence. Most of what is said online is written down and stays written down. This is an extraordinary gift and a trap. The gift is that the ethnographer can reconstruct a history they were not present for: the argument in 2019 that split the group is still readable, complete with the deleted-then-restored post everyone still talks about. The trap is that the archive tempts the researcher to substitute retrieval for participation. Reading five years of a forum in a week is not the same as having been there for five years, because you read it knowing how it turned out, in a compressed sitting, without the intervening ordinary days. The archive gives you the events and takes away the rhythm, and rhythm is much of what ethnography is for. The second is searchability and audience uncertainty. Because material is retrievable, its audience is not the audience present at the time of writing. danah boyd's account of networked publics identifies this as the structural condition of online sociality: the audience is invisible, contexts collapse, and public and private are not distinct states but a negotiated spectrum. For the researcher this has two consequences. Analytically, participants are writing for imagined audiences that include people not yet present, which shapes what they say in ways the text itself may not reveal; the interview is often the only way to find out who someone thought they were talking to. Ethically, the researcher is themselves one of those future audience members, and a quotation republished in a journal is a further collapse of context that the writer never consented to. The third is algorithmic mediation. What is visible on a platform is the output of a ranking system optimising for engagement, tuned continuously, and personalised to the viewer. Two researchers observing "the same" feed on the same day see different things. This breaks a basic ethnographic assumption — that what is in front of you is what is happening — and it has no clean fix. What it requires is documentation: record the account you observed from, its follow graph, its history, the date, and any known platform changes, and treat your observations as observations of one algorithmically constituted view rather than of the platform as such. Zeynep Tufekci's discussion of the "denominator problem" in social media research makes the general point sharply: without knowing the population from which visible material was drawn, proportions computed on it mean very little. Digital, virtual, networked: why the adjectives disagree Newcomers are often confused by the proliferation of names — virtual ethnography, netnography, digital ethnography, social media ethnography, online ethnography — and assume they are branding exercises. Mostly they are not; they encode different commitments. Virtual ethnography (Hine) came first and carried the early 2000s question of whether online interaction was real enough to study. Its answer was yes, and its method was to treat the internet as both culture and artefact. Netnography is Robert Kozinets's codified procedure, developed initially in consumer research and set out most fully in the successive editions of Netnography (2010; revised 2020). It is more prescriptive than most anthropological approaches — with defined stages of entrée, data collection, interpretation, and member checking — and it is candid about researching communities for applied and commercial purposes. That prescriptiveness is genuinely useful for students and for applied projects with timelines; it can also encourage a checklist relationship to fieldwork. Digital ethnography as used by Sarah Pink and colleagues in Digital Ethnography: Principles and Practice (2016) widens the frame again: the digital is not a separate site but part of everyday environments, studied alongside the sensory and material aspects of life. John Postill and Pink's earlier account of "social media ethnography" described the researcher's practice as catching up, sharing, exploring, interacting, and archiving — a messy set of daily routines rather than a residence. The practical upshot is not that one label is correct. It is that the labels come with different default answers to the boundary question, and choosing a label without choosing a boundary is how projects end up incoherent. Choosing a site you can actually study Three tests separate a workable site from an aspiration. Is there enough interaction to observe? A great many online spaces are archives of monologue: posts with no replies, comments with no responses to the comments. If your interest is in interaction — how members respond to a newcomer's disclosure, how a norm is enforced — then a site without threads cannot give it to you, however large it is. Before committing, read fifty randomly chosen items and count how many are parts of exchanges. This takes an afternoon and saves months. Is there a history you can reconstruct? Communities are constituted by their crises. If the platform deletes old content, or the group's history is in a channel you cannot access, you will be studying a permanent present, which makes it very difficult to tell a norm from a mood. Can you get lawful, documented access to it, at the depth the question requires? This is where many otherwise good designs die, and it is worth doing the feasibility check first rather than last. A question about private-group dynamics requires either participation with the group's knowledge or interviews; it cannot be answered from public data, and no amount of methodological ingenuity changes that. Chapter 5 deals with the access routes in detail. The point here is that site selection and access route are one decision, not two. A fourth consideration is less a test than a warning. Ease of access is not a reason to pick a site, and it is the single most common reason sites actually get picked. For roughly a decade, a disproportionate share of published social media research studied Twitter, for a reason that had nothing to do with Twitter's social importance: its API was open, its data was cheap, and its posts were short enough to model. Derek Ruths and Jürgen Pfeffer set out the resulting distortions in Science in 2014, and Tufekci made the related argument that Twitter had become the "model organism" of social media research — the fruit fly whose convenience shaped the questions the field asked. When academic access to that platform was withdrawn in 2023, a substantial body of accumulated method turned out to be tied to an interface rather than to a problem. The lesson is not to avoid popular platforms; it is to be able to say, in one sentence, why this site answers this question, and to have that sentence survive the platform changing its pricing. Where offline and online stop being separable An older debate asked whether online research needed offline "grounding" to be valid. That debate is over, and it ended in a draw that is worth understanding, because the draw is the working position most good projects now adopt. The position that online life is real life, studiable on its own terms, won the argument against dismissal. Nobody now seriously claims that a decade-long friendship maintained in a group chat is not a real friendship, or that norms enforced in a moderated forum are not real norms. But the position that online settings are separate worlds lost too, and lost harder. Nearly all the settings a contemporary researcher studies are continuous with the rest of participants' lives: the recovery forum's members go to clinics, the political community organises offline events, the fandom has jobs and rent. Miller's comparative work makes the case with unusual force, because it found that what social media did varied enormously by place — the same platforms functioned as conservative instruments of family surveillance in some settings and as instruments of escape in others. A study of the platform alone could not have seen that. The practical resolution is that the boundary of the field is set by the question, not by the technology. If the question is about how a moderation norm gets enforced, the norm's enforcement is observable in the forum and you may not need to leave it. If the question is about what membership in the group means to people, you will have to ask them, and what they tell you will be full of things you could not have seen. This has a consequence worth stating plainly, because it recurs in every subsequent chapter. Observation alone rarely answers questions about meaning. You can see what people do in a thread. You cannot see what they thought they were doing, who they thought was reading, what they decided not to post, or what happened in the direct messages that made the public exchange look the way it did. The invisible portion of online interaction is enormous, and it is systematically invisible — the most sensitive material moves to the most private channel. Designs that rely wholly on public traces are therefore designs that have accepted a specific, structured blindness. That can be a reasonable trade. It is not a reasonable thing to leave unmentioned. The platform is a participant, not a venue One more thing distinguishes a digital field, and it is easy to miss because it is everywhere: the site has an owner, and the owner is an actor in it. In a physical field, the built environment shapes interaction but does not intervene in it. Online, the environment writes rules, enforces them selectively, changes the interface without notice, ranks what members see, deletes accounts, and occasionally deletes the community. Any account of norms in an online group that ignores the platform's own governance has left out one of the three parties to every interaction — the poster, the audience, and the system that decides whether the post is seen. This shows up concretely in fieldwork. A sudden change in the tone of a forum may be a change in its membership, or it may be that a moderation tool was introduced. A drop in the visibility of a topic may reflect declining interest or a change in ranking. Members themselves theorise about this constantly — the vernacular vocabulary of "shadowbanning," "the algorithm," "getting reach," and deliberately misspelled words to evade automated detection is evidence that participants experience the platform as an agent with intentions. Those folk theories are excellent data in their own right, whether or not they are technically accurate, because they shape what people write and how they write it. There are two practical implications. The first is that a study should record the platform's governance context alongside its social one: the terms of service and community guidelines in force during the study period, notable policy or interface changes, and any moderation events the community discussed. Guidelines change often and are rarely archived by the platform, so save dated copies as you go. The second is that the researcher should resist inferring from absence. Content that is not visible may never have existed, may have been deleted by its author, may have been removed by moderators, or may simply not have been surfaced to the account you were watching from. These are analytically different, and only the community's own talk about removals usually tells you which occurred. Gabriella Coleman's ethnographies of free-software developers and of Anonymous are instructive here precisely because their subjects were unusually articulate about infrastructure. In communities that build and argue about the systems they inhabit, the platform's role is stated aloud. In most communities it is not, and the researcher has to learn to hear it anyway. What to write down before collecting anything A site definition that exists only in the researcher's head is a site definition that will drift. Write it down, in a page, before collection begins, and keep the page as a document to be revised with dates rather than replaced. It should state: what the field is, in one sentence, in terms a stranger would understand; what is inside it and what is adjacent but outside; the period covered and why those dates; the platforms and specific spaces included; how you will know if the boundary turns out to be wrong; and what you expect to be unable to see. That last item is the one people skip and the one reviewers notice. If the design is a following design in Hine's sense, the page instead records the rule for following — which connections you will pursue and which you will decline — because without such a rule, multi-sited research becomes tourism. None of this is bureaucracy. It is the difference between a study whose claims have a referent and a study that has collected a large amount of material about roughly a topic. The rest of this booklet assumes the page exists. The next chapter argues that it should be written alongside a second one, about harm, and that neither is finished until both are. Chapter 2: Ethics as Method, Not Paperwork In 2008 a research team released a dataset built from the Facebook profiles of an entire cohort of undergraduates at one American university. The data was anonymised by the standards of the day: names removed, the institution identified only as "an anonymous north-eastern American university." Within days, Michael Zimmer had established which university it was, using nothing more than the dataset's own documentation — the cohort size, the unusual range of majors, the study-abroad destinations. He published the analysis in Ethics and Information Technology in 2010 under a title that has since become a slogan: "But the data is already public." The case is worth keeping in mind for two reasons. The first is the obvious one about re-identification. The second is subtler and more important: nothing the researchers did was malicious, careless by the standards then current, or unapproved. They had ethics board clearance. The failure was a failure of method — of thinking about data collection, anonymisation, and release as separable technical steps rather than as one design problem with ethics running through it. That is the argument of this chapter. Ethics in online research is not a form to be filled in before the interesting work starts. It is a set of design decisions that determine what data you have, what context survives with it, what you can quote, and therefore what you can claim. Get it wrong early and you will find, at the writing stage, that you cannot publish the evidence for your best finding. Why "public" is not a permission The most persistent error in this field is the inference from publicly accessible to fair game. It is worth taking apart carefully, because it is not stupid — it is just wrong in a specific way. Accessibility is a property of a technical system. A post on an open forum can be retrieved by anyone with a browser. From this it does not follow that the poster consented to research use, that they understood the post's retrievability, or that republishing it causes no harm. Legal availability and ethical appropriateness are different questions, and the answer to the second cannot be read off the first. Nissenbaum's contextual integrity, introduced in Privacy in Context (2010), gives the better frame. Every context has norms governing the flow of information within it: what may be said, to whom, under what conditions it may be passed on. Privacy is preserved when flows respect those norms and violated when they do not — including when information that was already "public" moves into a context with different norms. A person who posts about relapse in a recovery forum has shared information in a context whose norm is that it stays among people who understand it. Quoting that post in a paper, which may then be indexed and surfaced to the person's employer via a search engine, is a flow across contexts. It may still be justifiable. It is not automatically permitted by the fact that the forum had no login. The Association of Internet Researchers has built its guidance on essentially this insight. The successive AoIR documents — the 2002 guidelines, Markham and Buchanan's 2012 revision, and Internet Research: Ethical Guidelines 3.0 (franzke, Bechmann, Zimmer and Ess, 2020) — decline to provide rules, on the deliberate ground that rules do not survive contact with new platforms. What they provide instead is a case-based, question-driven process: ask what the participants' reasonable expectations were, what harm could follow, how vulnerable the population is, and whether the research value justifies the flow. Reviewers in this field increasingly expect to see that reasoning shown, not a citation to a principle. A useful way to make the reasoning concrete is to place the material on a spectrum rather than in a box, as Table 2 does. Table 2. Settings on the accessibility–expectation spectrum, and what each implies. Setting Technically accessible to Members' likely expectation Re-identification risk Usual consent route Open forum, indexed by search engines Anyone Mixed; many assume obscurity High for verbatim quotes Notice to community; individual consent for quotes Large open platform, public posts Anyone Aware of visibility; rarely of research High if quoted; low if aggregated Aggregate freely; seek consent to quote Closed group requiring approval Approved members Confidentiality within group Very high Negotiated access with moderators plus individual consent Semi-public channel shared by link Anyone with the link Obscurity by default Very high Treat as private; consent required Direct or ephemeral messages Intended recipient Fully private Absolute Consent from all parties; usually solicited data only The table is a heuristic, not a licence. Two considerations override it. One is topic sensitivity: the same forum is a different ethical object when the thread is about kitchen renovation and when it is about suicidal ideation. The other is population vulnerability: minors, people in legal jeopardy, people whose identity is criminalised where they live, and people in acute distress warrant protection well beyond what the setting's accessibility suggests. Harm, specified "Risk of harm" is a phrase that invites hand-waving. It helps to name the mechanisms, because each has a different mitigation. Identification harm occurs when a research output lets someone be identified who did not expect to be. The main vector online is verbatim quotation: a distinctive sentence pasted into a search engine returns the original post, and from it the account, and from the account potentially a name, employer, and location. This is the single most under-appreciated risk in the field, and Chapter 9 deals with the response in detail. It is worth noticing now that it makes online research different from interview research in a way that cuts against normal practice — in interview studies, verbatim quotation is a mark of rigour; in studies of written online material, it can be a disclosure. Aggregation harm occurs when individually innocuous material is combined into a profile. A person's posting history across a decade, assembled into one file, reveals a great deal that no single post did. Researchers routinely assemble such files without thinking of them as dossiers. Group harm occurs when a community is characterised in ways that damage it collectively, even if no individual is identified. Studies of stigmatised communities frequently produce this, and it is not addressed by anonymisation at all. Naming the community is often the problem; several researchers now pseudonymise the forum as well as the members. Disruption harm occurs when the research alters the setting. An ethnographer who joins a small group and asks a lot of questions changes it; a researcher who posts a recruitment notice to a support forum imports a research frame into a space that people use to not be studied. Downstream harm occurs when data leaves the study. A dataset shared for replication, a corpus deposited in a repository, a set of quotes reproduced in a press release — each is a further flow. Against this list, the researcher's commonest reassurance — "I will anonymise" — is a response to one mechanism out of five. Consent when consent is impractical The hardest practical question is what to do when the population is too large, too transient, or too anonymous to consent in the conventional sense. There is no single answer, but there is a defensible ordering of options, and the discipline is to work down it rather than jumping to the bottom. Individual informed consent remains the default where interaction is involved: interviews, participation, anything solicited. It is also achievable more often than people assume for quotation. Contacting an account to ask permission to quote a specific post works reasonably often, particularly if the request is specific and shows the intended text. Community-level consent — negotiating with moderators or administrators, posting a notice, allowing an opt-out period — suits participant observation in bounded groups. It is weaker than individual consent because moderators do not speak for members, but it satisfies the norm that a community should not be studied in secret. It also tends to improve the research: members who know a researcher is present will often tell them things. Waiver with mitigation is appropriate for large-scale study of public material where contacting individuals is genuinely impossible and the topic is not sensitive. Here the obligation shifts entirely onto mitigation: no verbatim quotation of identifiable material, no republication of usernames, aggregate reporting, and a paraphrase protocol for illustrative material. Retrospective or member-checking approaches, in which findings are shown to community members before publication, do not substitute for consent but frequently catch errors of interpretation and unanticipated harms. What is not on the list is the reasoning that because consent is hard, it does not apply. The size of a dataset has no bearing on whether the people in it are people. There is a specific hard case worth naming: research on communities the researcher considers harmful — extremist movements, harassment networks, organised disinformation. Here the consent norms genuinely bend, because informing a group that they are being studied may endanger the researcher and will certainly change what is said. Covert observation of such groups has a long and respectable tradition in offline sociology. The bending is not unlimited, though: the researcher still owes non-identification to individuals who are not public figures, still owes accuracy, and still has to have submitted the design for review rather than deciding alone. Researcher safety planning belongs in the same conversation, since studying such groups carries a real risk of retaliation. Law, regulation, and terms of service Ethics and law are separate, and confusion between them causes trouble in both directions — researchers who assume that lawful means permissible, and researchers who assume that any terms-of-service question makes a project impossible. Data protection law applies to most of this work. Under the EU General Data Protection Regulation, online posts about identifiable people are personal data whether or not they are public, and some — health, political opinion, sexual orientation, religion — fall into the special categories that require additional justification. GDPR does contain provisions that accommodate research: Article 89 permits derogations for scientific research subject to safeguards, and a public-interest or legitimate-interests basis is often available where consent is not. The practical effect is not that research is exempt but that it must be documented — purpose, lawful basis, minimisation, retention period, security. UK and EU institutions will generally require a data protection impact assessment for a project of any scale. Note that the regulation applies to pseudonymised data, since pseudonymised data remains personal data; only genuinely anonymous data falls outside, and true anonymisation of text is very difficult. Human-subjects regulation in the United States takes a different route, through the Common Rule and institutional review boards. The definition of human-subjects research turns on interaction or identifiable private information, which leaves a substantial category of public online data outside formal IRB jurisdiction. This exemption is a jurisdictional fact, not an ethical verdict; the Menlo Report (2012), which adapted the Belmont principles for information and communications technology research, was written partly because so much consequential work fell outside the existing framework. Terms of service are the most misunderstood layer. Platform terms typically restrict automated collection, bulk storage, and redistribution. Breaching them is a contract matter, and the main practical consequences are account termination, loss of access, and institutional exposure; several journals and funders now ask directly whether data was collected in compliance. In United States law the question of whether breaching terms also constitutes computer misuse has narrowed considerably: Van Buren v. United States (2021) read the Computer Fraud and Abuse Act's "exceeds authorized access" clause narrowly, and the hiQ Labs v. LinkedIn litigation in the Ninth Circuit held that scraping publicly available data was unlikely to constitute unauthorised access under the statute — while leaving contract and other claims alive. In Sandvig v. Barr (2020) a district court held that mere terms-of-service violations in the course of research did not constitute a criminal CFAA offence. The net position is that the criminal risk to an academic researcher collecting public data is much lower than the folklore suggests, and the contractual, institutional, and ethical constraints are entirely intact. This booklet takes a firm line on what follows. Working within an interface a platform provides, on terms it publishes, is research. Defeating technical access controls, using credentials that are not yours, creating accounts to gain entry to private spaces under false pretences without review, or attempting to re-identify individuals in a de-identified dataset are not methods and are not treated as such here. The reason is not timidity. It is that each of those acts destroys the thing the research depends on — the possibility of saying honestly, in the write-up, how the material was obtained. The regulated alternative is worth knowing about, because it is new. The EU's Digital Services Act creates, in Article 40, a mechanism by which vetted researchers can apply through national Digital Services Coordinators for access to data from very large platforms for research on systemic risks. The Commission adopted the delegated act setting out the operational rules in July 2025, and the regime has been coming into effect since. It is narrower than its advocates hoped — it is oriented to systemic risk rather than to social research generally, and it is institutionally heavy — but it is the first legal route to platform data that does not depend on a company's goodwill. Three cases worth having in mind Abstract principles are easier to apply when attached to episodes that actually happened, and this field has a small canon of instructive failures. The emotional contagion experiment reported by Adam Kramer, Jamie Guillory and Jeffrey Hancock in the Proceedings of the National Academy of Sciences in 2014 manipulated the emotional valence of content shown in hundreds of thousands of Facebook news feeds and measured the effect on users' own posting. The science was modest; the reaction was not. The core objection was that experimental manipulation of mood without consent is not made acceptable by being routine product optimisation, and the case established in most researchers' minds a distinction between what a platform may do to its users commercially and what a researcher may do to them scientifically. The journal appended an editorial expression of concern. The durable lesson concerns collaboration: when academic researchers work with platform data scientists, the ethical regime that applies is the academic one, and it has to be established before the experiment, not defended after it. The OkCupid dataset released in 2016 by a researcher who posted profile data on tens of thousands of users to an open repository — including usernames, ages, locations, and answers to intimate questionnaire items — is the purest available illustration of the "already public" fallacy. The data had been accessible to logged-in users of the service. Compiling it into a downloadable file changed what it was, from scattered profiles readable one at a time by people with a reason to look, into a searchable dossier. The repository removed it. No anonymisation was attempted, and the defence offered was that none was needed because nothing had been hidden. The hacked-data question arises whenever a breach dumps a corpus that would answer a research question: the Ashley Madison leak in 2015 is the standard example, and there have been many since. The material is undeniably available and undeniably informative. The arguments against using it are that the data subjects suffered a harm the researcher would be extending, that the research would create an incentive structure rewarding breaches, and that provenance cannot be verified. The narrow exceptions people defend — research on the breach itself, or on named public actors where the public interest is substantial — prove the rule by how narrow they are. The general position is that a dataset's origin is part of its ethics, and that "someone else did the wrong" is not a laundering mechanism. What the three have in common is that each involved competent researchers who had asked whether the data was available rather than whether the flow was appropriate. Building the ethics into the design Practically, this chapter's argument reduces to doing four things at outline stage rather than at submission stage. Decide the consent route before choosing the site, because the route constrains which sites are feasible. Decide the quotation policy before collecting, because it determines whether you need to be able to contact posters. Decide the retention and storage plan before the first file exists, because retrofitting security to a folder of scraped posts on a laptop is not really possible. And write down, in advance, what you will do if you encounter disclosure of imminent harm — this happens more often in online fieldwork than researchers expect, and deciding in the moment goes badly. A practice that helps with all four is to keep an ethics log beside the fieldnotes: a dated record of decisions, the reasoning behind them, and the things that unsettled you. It costs ten minutes a week. It produces, at the end, both the material for an honest methods section and the evidence that the judgement calls were judgement calls rather than defaults. Chapter 3: Being There When There Is No There The classical justification for participant observation is that some things can only be learned by being present while they happen. You cannot get the texture of a workplace from interviews alone, because people describe their jobs in terms of official purposes and forget the improvisations that actually make the place run. The ethnographer's presence over time produces knowledge of a specific kind: tacit, comparative, and grounded in having seen enough ordinary instances to recognise an unusual one. Online, the difficulty is not access — you can read the forum this afternoon — but presence. Access is trivially available and presence is genuinely hard, which is the reverse of the offline situation and the reason that much published "online ethnography" is not ethnography at all but a content analysis with a longer reading period. This chapter is about how presence is constructed when there is nowhere to stand. The lurking question Every online fieldworker begins as a reader and has to decide how long to stay one. Reading without posting is not a moral failing. It is, in most online spaces, the majority practice; the great bulk of any community's membership never posts, and the researcher who only reads is behaving like most members. In spaces where reading is the normal mode, a period of sustained reading is genuine immersion rather than a substitute for it. But reading without posting has two specific limits. First, it does not test understanding. The ethnographer who participates finds out when they have got something wrong, because their contribution lands badly; the ethnographer who only reads can hold an incorrect interpretation indefinitely with nothing to correct it. Second, undisclosed reading forecloses the relationships that yield interviews, clarifications, and the eventual member checks that protect against misreading. There is also an ethical dimension. Observing a group at length without telling them is covert research, whatever the platform's accessibility. In an open, unbounded, high-traffic space this is usually unproblematic, since there is no group to inform. In a small, bounded community with a membership that would recognise a stranger's presence, it is a decision that needs justifying rather than a default. A workable default: read widely and without announcement while scoping, because scoping is not yet the study; disclose before the study proper begins if the community is bounded enough for disclosure to mean something; and, once disclosed, participate as a member would, at whatever level is honest for someone in your position. What that level is varies enormously, and Table 3 is not needed to say so — the honest description is that you should be a recognisable kind of participant rather than a category-defying one. Disclosure as a practical problem Telling a community you are studying it is easy to say and awkward to do. Three approaches recur. The moderator route is the most common for bounded groups: approach the moderators or administrators privately, explain the project, ask permission, and agree what the membership will be told. Its advantage is that moderators know their community and will tell you things about it in the negotiation itself — often the single most informative conversation of the early fieldwork. Its risk is capture: a study conducted with the moderators' blessing can quietly become a study conducted from the moderators' point of view, and moderators are usually in conflict with some part of the membership. The profile route works in spaces with persistent identity: the researcher's account states plainly who they are and what they are doing, and their posts are visible as coming from someone doing research. This is passive disclosure. It works where profiles are read and fails where they are not. The periodic announcement route suits long studies in high-turnover spaces: a post at intervals restating the project, with an email address and an opt-out. It is the only approach that deals with the fact that a community's membership in month twelve is not the membership that was informed in month one. None of these achieves informed consent from everyone observed, and the honest methods section says so. What they achieve is that the research is not secret, which is a lower but real standard. Expect disclosure to change the setting, and treat the change as data. Communities respond to being studied in patterned ways — a flurry of meta-discussion, some hostility, some enthusiastic volunteering, some members going quiet. The hostility is often the most informative: what people fear a researcher will do with their words tells you about the community's experience of exposure, which is frequently the very topic under study. Constructing presence Presence is made of three things that come free offline and have to be built online: rhythm, co-presence, and accountability. Rhythm means being there on the community's schedule rather than your own. Every online space has a temporal structure — the forum where substantive posts appear on weekday evenings, the chat server that is dead until a European afternoon, the group whose real activity happens in a two-hour window after a weekly broadcast. A researcher who reads in one long weekly session experiences the space as an archive. A researcher who checks in daily at the right times experiences it as a place where things are happening now, with the attendant sense of waiting, missing things, and catching up. Postill and Pink's description of social media ethnography as a set of daily routines — catching up, sharing, exploring, interacting, archiving — is a description of rhythm construction. Practically this means setting a schedule and keeping it for the duration, including through the dull weeks. The dull weeks are where the baseline comes from. Co-presence means being in the space at the same time as others, which written asynchronous platforms make oddly difficult. The methods that generate it are live ones: voice channels, streams, scheduled events, synchronous chat, games. Boellstorff's virtual-worlds fieldwork was possible partly because avatars produce a strong sense of shared presence; text forums produce almost none. Where a community has any synchronous component, it usually repays disproportionate attention, because it is where the relationships that sustain the asynchronous parts are made. Accountability means being someone the community can hold responsible. It is produced by the small, cumulative business of using a consistent identity, answering when addressed, acknowledging correction, following the norms including the trivial ones, and not disappearing. It is also produced by restraint: not asking too many questions too early, not treating every exchange as an elicitation, not being the member who only ever appears when something interesting happens. There is a version of online ethnography that consists of joining forty spaces and being present in none. It produces breadth of a kind. It does not produce the thing participant observation is for. Fieldnotes without the drive home Offline ethnographers have a structural advantage they rarely notice: the journey home. The interval between the field and the desk is when raw experience gets turned into something writable. Online there is no interval. The field is the same screen as the notes, and the temptation is to skip the notes entirely because the material is already saved. That temptation should be resisted absolutely, because saved material is not fieldnotes. Robert Emerson, Rachel Fretz and Linda Shaw's Writing Ethnographic Fieldnotes (1995) makes the point that fieldnotes are already interpretation — selection, description, and the first analytic moves — and that the writing is where the ethnography happens. An archive of posts contains none of that. It contains what was said, not what it was like, what you noticed, what surprised you, or what you failed to understand. A workable online fieldnote practice has four layers, kept separate. The log is a dated record of what you did: which spaces you visited, for how long, what you posted, what you collected. It is dull and it is what makes the eventual methods section true. The observational note is written immediately after each session, in prose, describing what happened as though to someone who was not there. The discipline is to describe before interpreting: who spoke, what the exchange was about, how it went, what the tone was. Twenty minutes of this after a session is the core of the practice. The analytic memo is where interpretation is allowed. Written less often — weekly is usually right — it develops emerging ideas, notes patterns, and, crucially, records when a previous interpretation has been overturned. The overturning is the evidence that the fieldwork is doing something. The reflexive note records the researcher's own position and reactions: irritation at a member, boredom, the pull of taking sides in a conflict, the feeling of being an intruder. This is not self-indulgence. In online fieldwork particularly, the researcher's reactions are the main instrument for detecting the emotional texture of a space, and unexamined reactions leak into analysis as unexplained judgement. Screenshots and saved threads attach to the observational note as exhibits, with the date and the account they were viewed from. Save the context, not just the item: a reply with the post it replies to, a comment with the thread around it. Decontextualised captures are almost impossible to interpret six months later, and six months later is when you will be writing. One specific discipline is worth adopting: note what you did not understand. Online spaces are dense with in-jokes, abbreviations, references to prior events, and deliberately oblique phrasing. The impulse is to look up the unfamiliar and move on. The better practice is to record the incomprehension, because the list of things that were once opaque and later became obvious is a precise record of what you learned, and it is the material from which you will later explain the community to readers who start where you started. Participation without distortion The reflexive worry of all participant observation — that the observer changes what is observed — takes a particular form online, because contributions are permanent, public, and quotable. An offline ethnographer's remark in a meeting is gone in a moment. A researcher's post in a forum is a durable artefact that will be read by people who join later, may be quoted back, may be screenshotted into another context, and constitutes a standing position on whatever it addressed. The asymmetry argues for a more conservative participation style than one would adopt offline: contribute genuinely, but be aware that every contribution is also a publication. Two specific traps are worth naming. The first is elicitation by post. Asking the group a question is fast and produces material, and it is very tempting. But a question posted by a known researcher reframes the space as a research site for as long as the thread lives, and it produces answers of a particular kind — considered, performed, addressed to the researcher and to the audience simultaneously. Used sparingly and with awareness, this is a legitimate method. Used as the main mode, it turns ethnography into a slow, public, badly controlled survey. The second is the standing that comes from attention. Researchers are attentive, articulate, interested in what people say, and generous with replies. In many online communities this is rare, and it earns status quickly. A researcher can find themselves, within months, an influential member of the community they are supposed to be observing. This is not always avoidable and is not always bad — it produces access — but it needs recording in the reflexive notes and reporting in the write-up. Following across platforms Few communities live in one place any more. A group with a public forum will have a chat server, some members will have a shared channel on a messaging app, and a subset will be connected on a general-purpose platform where they also conduct the rest of their lives. The fragmentation is not incidental; it is how people manage audience. Material that is fit for the public forum goes there, material for the inner group goes to the chat, and the genuinely sensitive goes to direct messages. This produces the practical problem of how far to follow. The maximal answer — follow everywhere the members go — is unworkable and often unethical, since consent obtained in one space does not extend to another and since the more private spaces are private for reasons. The minimal answer — stay in the one space you negotiated access to — risks studying the community's front stage and reporting it as the whole performance, in the way Goffman's distinction between front and back regions in The Presentation of Self in Everyday Life (1959) would predict. The defensible middle is to follow deliberately and to document the rule. Decide in advance which spaces are in scope; negotiate access to each separately rather than assuming transitivity; and when a space is out of scope, say so in the write-up and say what you therefore cannot claim. If members tell you about what happens in the private channel, that is reported speech and can be treated as interview material, with the usual caveats about how people describe their own backstage. There is a related temptation to resist: assembling profiles of individual members by following them across platforms. Even where every element is public, the compiled picture is a dossier, and building one about a research participant without their knowledge is closer to investigation than to ethnography. Where biographical depth on particular members is genuinely needed, the route is to ask them. Leaving the field Online fieldwork has no natural end. There is no flight home, no last day, no ceremony. Studies therefore tend to trail off, with the researcher gradually reading less until they stop, which is both analytically sloppy and, in a community where relationships were formed, a small betrayal. Set an end date and observe it. Announce departure in the spaces where you disclosed presence, thank people, say what will happen next and roughly when, and give a route by which members can reach you about the eventual output. If you promised member checking, do it before the announcement rather than after, while people still recognise your name. Decide also what happens to your account and your contributions. Deleting a research account removes your own posts from threads and can make other people's contributions unintelligible, so deletion is rarely the right answer; leaving the account dormant with a profile note explaining the project's status usually is. If the community's norms expect departing members to say goodbye, say goodbye in the form the community uses, not in the form an institution would draft. Two further obligations tend to be forgotten. The first is to the moderators who granted access: send them the eventual output, or a readable summary of it, before it appears rather than after. The second is to yourself, in the form of a final reflexive memo written at the point of exit, describing how the community looked to you at the end and how that differed from the beginning. Written a year later, when the analysis is finished, that memo will be the only surviving record of what you actually believed while you were in the field, as opposed to what you concluded afterwards. Knowing when observation has run out Participant observation answers questions about practice: what people do, how interaction is organised, what the norms are and how they are enforced, what counts as a competent contribution, how newcomers are inducted, how conflicts unfold. It does not answer questions about meaning, motivation, or the unobserved. It cannot tell you why someone posted, what they considered and discarded, what the private channel said, how the group figures in the rest of their life, or what they think the researcher wants to hear. Online, the unobserved portion is unusually large, because the platform architecture reliably sends the most consequential talk to the least visible channel. The practical consequence is that virtually every serious online ethnography is a mixed design, in which observation establishes what happens and interviews establish what it means to the people it happens to. Kozinets's netnographic procedure formalises this, and most anthropologically inclined work arrives at it informally. Knowing when to move is a matter of judgement, but there is a reliable signal: you have been observing long enough when you can predict, with reasonable accuracy, how the community will respond to a new event — a newcomer's first post, a policy change, a member's departure — and when the predictions start being right. At that point further observation yields diminishing returns, and the questions that remain are questions only participants can answer. The next chapter is about how to ask them, in a medium where the conversation may take three weeks. Hashtags: #QualitativeInquiryInDigitalSpace #DigitalEthnography #OnlineEthnography #Netnography #VirtualEthnography #DigitalFieldwork #ParticipantObservation #AsynchronousInterviewing #OnlineDiscourseAnalysis #ComputationalTextAnalysis #SentimentAnalysis #TopicModeling #NetworkMining #SocialNetworkAnalysis #ContextualIntegrity #DigitalResearchEthics #OnlineConsent #ReIdentificationRisk #PlatformGovernance #AlgorithmicMediation #DataDonation #PlatformAPIs #DigitalServicesAct #MixedMethodsDigitalResearch #FutureOfDigitalQualitativeInquiry

  • Proteomics by Design (Tandem Mass Tagging, DIA, and Quantitative Bioinformatics)

    Download the Book (PDF): Introduction A proteomics experiment never measures proteins. It measures ions: charged fragments of peptides, which are themselves fragments of proteins, flying through electric fields in an instrument that sees only mass-to-charge ratios and arrival times. Everything a biologist eventually reads in a results table, whether a list of eight thousand quantified proteins, a volcano plot of changes after drug treatment, or a claim that a kinase is phosphorylated at a particular serine, is the output of a long chain of inference that starts with those ions and works backward to biology. Each link in the chain was chosen by someone. Most of the choices were made before the first sample was ever loaded. That is the argument of this book. Quantitative proteomics is a design discipline before it is a measurement discipline. The depth of a proteome, the precision of a fold change, the trustworthiness of an identification and the honesty of a p-value are not properties of the mass spectrometer. They are properties of an entire workflow, from the buffer used to lyse the cells to the statistical model that tests for differential abundance. The two acquisition strategies that dominate contemporary discovery work, isobaric labelling with tandem mass tags (TMT) and data-independent acquisition (DIA), are best understood not as rival technologies but as two coherent sets of design decisions, each of which buys certain strengths by accepting certain costs. Choosing between them well requires understanding where those costs come from. The problem proteomics is trying to solve The human genome encodes roughly twenty thousand proteins. That number is deceptively small. Alternative splicing, proteolytic processing and hundreds of kinds of post-translational modification multiply it into a vast population of distinct molecular species, usually called proteoforms. More troublesome than the number is the range of abundances. In a typical cultured human cell line the most abundant proteins, such as histones, ribosomal proteins, cytoskeletal proteins and glycolytic enzymes, are present at millions of copies per cell, while transcription factors and many signalling proteins may be present at a few hundred or a few thousand. In blood plasma the problem is far worse: a handful of proteins, led by albumin, make up most of the protein mass, while the cytokines and tissue leakage markers of greatest clinical interest sit many orders of magnitude lower. Leigh Anderson and Norman Anderson's widely cited survey of the plasma proteome described that span as exceeding ten orders of magnitude. No analytical instrument covers such a range in a single measurement. A mass spectrometer has a finite capacity for ions at any instant, a finite number of spectra it can record per second, and a detection limit below which a signal disappears into noise. Every proteomics workflow is therefore an exercise in allocating scarce resources: separation capacity, instrument time, ion injection time and, eventually, statistical power. The practical question a researcher faces is never "how do I measure the proteome?" but "given what I want to know, how should I spend what I have?" Why bottom-up This book concerns bottom-up proteomics, sometimes called shotgun proteomics, in which proteins are enzymatically digested into peptides before analysis. The alternative, top-down proteomics, analyses intact proteins and preserves information about which modifications occur together on the same molecule. Top-down methods have advanced considerably, but for deep, quantitative interrogation of complex proteomes the bottom-up approach remains dominant, for reasons that are chemical rather than fashionable. Peptides of roughly seven to thirty amino acids are soluble, separate well by reversed-phase chromatography, ionise efficiently by electrospray, and fragment predictably in the gas phase. Their masses fall in a range where instruments are fast and sensitive. Proteins, by contrast, are heterogeneous in size, solubility and charge, and large ones fragment poorly. The price of digestion is the loss of the direct link between a measured peptide and the protein it came from. A peptide shared between two related proteins, or two isoforms of the same gene, cannot by itself say which of them was present. Reconstructing proteins from peptides, a problem known as protein inference, is one of the recurring difficulties the later chapters return to, because it affects identification, quantification and the calculation of error rates alike. Two philosophies of acquisition For most of the history of the field, discovery proteomics relied on data-dependent acquisition (DDA). The instrument surveys the peptides eluting from the chromatography column at a given moment, picks the most intense few, isolates each in turn, fragments it and records the fragment spectrum. The approach is intuitive and produces clean spectra that link one precursor to one set of fragments. Its weakness is that the choice of which peptides to fragment is partly random, driven by intensity and timing, so that the same peptide may be picked in one run and missed in the next. Across a study of many samples the result is a patchwork of missing values. The two strategies at the centre of this book represent two different answers to that weakness. Isobaric labelling, commercialised as iTRAQ and TMT in the early 2000s, attaches chemical tags to the peptides of several samples, pools them and analyses them together. Because the tags have identical total mass, the same peptide from every sample appears as a single precursor; only upon fragmentation do the tags release small reporter ions of distinct masses whose intensities reveal the relative amounts in each sample. If a peptide is picked for fragmentation, it is quantified in all samples at once. Missing values within a multiplexed set largely disappear. The costs are reagent expense, a ceiling on the number of samples per set, and a distortion called ratio compression, in which co-isolated peptides contaminate the reporter signal and pull measured ratios toward one. Data-independent acquisition takes the opposite route. Instead of choosing peptides, the instrument systematically fragments everything within a series of predefined mass windows, cycling through them fast enough to sample every eluting peptide several times across its chromatographic peak. The result is a complete, reproducible record of the fragment ions produced by every sample, at the cost of spectra that are highly multiplexed and therefore hard to interpret. DIA moved from a niche technique to a mainstream one in the decade after the SWATH-MS method was published in 2012, as faster instruments and new software, much of it based on machine learning, learned to untangle those spectra. Neither approach is universally superior. Controlled comparisons, including a 2019 study from Biognosys and the Leibniz Institute on Aging that matched instrument time between a TMT SPS-MS3 workflow and a single-shot DIA workflow, have found the familiar pattern: isobaric labelling identified somewhat more proteins and gave slightly better precision in that design, while DIA gave better accuracy of fold changes, and both detected true changes at similar rates. The balance shifts with sample number, sample amount, instrument generation, the biological question and the budget. A researcher who understands why each method behaves as it does can make that decision deliberately rather than by habit or by what the local core facility happens to run. Why error rates deserve a whole chapter The final third of the workflow, the bioinformatics, is where the most consequential and least visible decisions are made. A modern instrument can record tens of thousands of spectra per hour. Matching them to peptide sequences is a search problem with an enormous space of candidates, and some fraction of the matches will be wrong. The field's solution, the target-decoy strategy for estimating the false discovery rate (FDR), is one of the genuine intellectual achievements of computational proteomics, and it is widely misunderstood. A one percent FDR at the level of peptide-spectrum matches does not imply a one percent FDR at the level of proteins; error rates that are controlled in each run separately compound when the runs are merged; and a 2025 study using so-called entrapment experiments reported that none of the DIA search tools it examined consistently controlled the peptide-level FDR, with particularly poor performance on single-cell datasets. These are not esoteric concerns. They determine whether a reported protein list can be trusted. How the book is organised The chapters follow the path of a sample. Chapter 1 begins with sample preparation and digestion, the step where more experiments fail than at any other and where the bottom-up bargain is struck. Chapter 2 treats separation, both the online liquid chromatography that feeds the instrument and the offline fractionation that multiplies depth, as a resource to be budgeted. Chapter 3 explains what happens inside the mass spectrometer, how data-dependent acquisition works and why its stochastic sampling shaped everything that came after. Chapter 4 is devoted to isobaric labelling: the chemistry of TMT, the problem of ratio compression and the instrumental remedies for it, and the design of multi-batch studies. Chapter 5 does the same for DIA: window schemes, spectral libraries and library-free analysis, ion mobility, and the newest narrow-window methods. Chapter 6 turns to identification itself, from classical database search to predicted spectra and machine-learned rescoring. Chapter 7 is about counting errors, the logic of false discovery rates at every level from spectrum to protein. Chapter 8 follows the numbers from peptides to biological conclusions: normalisation, missing values, statistical testing and the question of how to choose between TMT and DIA for a given study. The conclusion argues what these threads imply for how proteomics experiments should be planned and judged. The reader assumed throughout is a working scientist from an adjacent field: a cell biologist who sends samples to a core facility, a geneticist who wants to integrate protein data with transcriptomes, a statistician asked to analyse a proteomics dataset, or a clinician evaluating a biomarker paper. No prior training in mass spectrometry is required, but the book does not simplify the substance. Where the field has settled a question, it says so. Where the question is open, it says that too. Chapter 1: From Tissue to Peptides The most expensive mass spectrometer in the world cannot recover a protein that was never extracted, a peptide that was never cut, or a sample that was contaminated with polyethylene glycol from a plastic tube. Experienced proteomics facilities tell a consistent story about failed projects: the cause is rarely the instrument and rarely the software. It is usually the sample. Sample preparation is where the bottom-up bargain is struck, where the proteome is converted into a peptide mixture whose composition will constrain every downstream measurement. It deserves to be treated as part of the experimental design rather than as a bench chore that precedes it. The goal of sample preparation can be stated simply. Starting from cells, tissue, biofluid or a purified complex, it should produce a clean mixture of peptides that represents the original proteins as completely and as reproducibly as possible, in a solvent compatible with reversed-phase chromatography and electrospray ionisation, and free of anything that suppresses ionisation or fouls the instrument. Every clause of that sentence hides a trade-off. Extraction, denaturation and the chemistry of artefacts Proteins must first be released from their biological matrix and unfolded so that a protease can reach its cleavage sites. The most effective way to solubilise membrane proteins, chromatin-bound proteins and aggregates is to use strong ionic detergents, above all sodium dodecyl sulfate (SDS), often combined with heat and mechanical disruption such as sonication or bead beating. SDS is also poison for mass spectrometry. It binds to reversed-phase columns, suppresses ionisation of co-eluting peptides and produces intense background ions. Chaotropes such as urea and guanidinium chloride denature proteins without the same chromatographic problems but are less effective at dissolving membranes, and urea brings a chemical hazard of its own: in solution it slowly decomposes to cyanate, which reacts with amine groups on proteins to form carbamylation artefacts. A urea buffer left warm overnight can modify a substantial fraction of lysine residues and peptide amino termini, shifting their masses and, in isobaric labelling workflows, blocking the very amines that the tags must react with. Much of the methodological history of proteomic sample preparation since 2009 is a history of ways to use detergents for extraction and then remove them before analysis. Three families of method now dominate, and each is worth understanding because each carries characteristic strengths. Filter-aided sample preparation, published by Jacek Wiśniewski, Matthias Mann and colleagues in Nature Methods in 2009, lyses the sample in SDS and then loads it onto a centrifugal ultrafiltration unit. The detergent is exchanged away by repeated washes with urea, the proteins remain on the filter, and digestion takes place on the membrane, after which the peptides are spun through. The method made SDS lysis compatible with shotgun proteomics and was widely adopted, though it is slow, involves many centrifugation steps and can lose material on the filter, especially at low input. Suspension trapping, commercialised as S-Trap and first described by Alexandre Zougman and colleagues in 2014, also begins with SDS lysis. Adding phosphoric acid and a high concentration of methanol causes the proteins to form a fine particulate suspension, which is trapped in a quartz or silica filter plug in a pipette tip or plate; the SDS is washed through, and digestion takes place in the trap. It is fast and handles large and difficult samples well. Single-pot, solid-phase-enhanced sample preparation, or SP3, introduced by Christopher Hughes, Jeroen Krijgsveld and colleagues in 2014 and refined in a 2019 protocol, uses carboxylate-coated paramagnetic beads. In the presence of a high concentration of organic solvent such as ethanol or acetonitrile, proteins aggregate onto the bead surface regardless of detergent. The beads are held by a magnet, the contaminants are washed away, and digestion proceeds with the proteins still associated with the beads. Later work showed the mechanism is better described as solvent-induced protein aggregation captured by the beads than as specific binding, which explains why the method tolerates almost any lysis buffer. Because SP3 works in a single tube or well and needs no centrifugation, it is readily automated on liquid-handling robots and scales down to low microgram and sub-microgram inputs, which is why it has become a default in many laboratories running large sample series. No one method is best for every sample. For formalin-fixed paraffin-embedded tissue, for example, the cross-links introduced by fixation must be reversed by prolonged heating before any of these methods can recover proteins efficiently; for plasma, the challenge is not extraction at all but dynamic range. What matters for design is consistency. Within a single study every sample should go through the same protocol, ideally in the same batch or in batches balanced across experimental groups, because differences in extraction efficiency between batches will otherwise masquerade as biology. Before digestion, disulfide bonds between cysteine residues are broken with a reducing agent, commonly dithiothreitol or tris(2-carboxyethyl)phosphine, and the freed thiols are capped with an alkylating agent so that they cannot re-form bonds. Iodoacetamide is the classical choice; it adds a carbamidomethyl group to cysteine, increasing its mass by about 57 daltons, and search software is told to expect that modification on every cysteine. Chloroacetamide reacts more slowly and produces fewer side reactions at non-cysteine residues, which is why several groups adopted it, but it was subsequently found to increase oxidation of methionine. These details sound minor. They are not. Every unexpected chemical modification introduced during preparation splits the signal of a peptide into several species of different mass, only one of which the search engine is looking for, and so reduces both identification rates and quantitative precision. Over-alkylation of lysines or peptide N-termini with iodoacetamide can even mimic the mass of biological modifications: a double carbamidomethylation of lysine has almost the same mass as the diglycine remnant that ubiquitination leaves behind after tryptic digestion, a coincidence that has caused false ubiquitination site assignments. The broader lesson is that sample preparation chemistry writes directly into the data. A search engine can only find what it is told to look for, and the unexplained spectra in any dataset are partly the fingerprints of chemistry that nobody accounted for. Digestion and clean-up Trypsin is the workhorse enzyme of bottom-up proteomics, and the reasons are instructive. It cleaves on the carboxyl side of lysine and arginine, residues that occur often enough in proteins that most tryptic peptides fall in a length range well suited to chromatography and fragmentation. Each tryptic peptide ends in a basic residue, which carries a positive charge; combined with the free amino terminus, this gives most tryptic peptides a charge of two or three in electrospray, and the basic C-terminus directs fragmentation to produce a clean series of y-type fragment ions that is easy to interpret. Trypsin is cheap, highly specific and active in the presence of moderate concentrations of denaturants. The textbook rule adds that trypsin does not cleave when the lysine or arginine is followed by proline. In practice, large-scale data have shown that cleavage before proline does occur at a meaningful rate, and modern search engines generally ignore the proline rule. More consequential are missed cleavages. Trypsin cleaves inefficiently at sites next to acidic residues, at sites with consecutive basic residues, and at lysines in structured regions. Missed cleavages generate longer peptides containing an internal lysine or arginine, splitting the signal of a region of the protein among several peptide species. A common remedy is to digest first with endoproteinase Lys-C, which cleaves after lysine, tolerates high urea concentrations and is more complete at lysine sites, and then with trypsin. Search settings typically allow up to two missed cleavages; a sudden rise in the missed-cleavage rate across a batch is one of the most useful quality-control signals that digestion went wrong. Trypsin also has blind spots. Regions of a protein that are very rich in lysine and arginine produce peptides too short to be informative, while long stretches without them produce peptides too long to detect well. Sequence coverage of a protein from a tryptic digest is therefore rarely complete, often well below half, and certain modification sites are systematically invisible. Alternative proteases, including chymotrypsin, Glu-C, Asp-N, Lys-N and alpha-lytic protease, produce overlapping but different peptide sets, and combining several digests raises coverage substantially. That is valuable when the goal is to map every modification site on a single protein or to identify sequence variants, but it multiplies instrument time, so for deep quantitative profiling of many samples trypsin alone remains the standard. Digestion is also where quantitative bias enters. Peptides are the unit of measurement, but the efficiency with which each peptide is released from its protein depends on local sequence and structure. That efficiency is constant from sample to sample only if the digestion conditions are constant. When they are not, peptide-level differences appear that have nothing to do with protein abundance, and they are indistinguishable from real changes unless the study design balances them. After digestion, peptides are desalted, usually on a reversed-phase C18 material, to remove salts, residual reagents and buffer components. The stop-and-go extraction tip, or StageTip, described by Juri Rappsilber, Yasushi Ishihama and Matthias Mann in 2003, packs small discs of C18 membrane into pipette tips; it is cheap, easily parallelised and still widely used. Plate-based cartridges and automated systems do the same job at scale. Peptide yield should then be measured, because loading a consistent amount of material onto the column is a precondition for reproducible quantification, and because the amount of peptide available sets hard limits on later choices such as how many fractions can be made and whether isobaric labelling, which requires a certain amount of material per channel, is practical. Contamination deserves specific mention because its effects are so disproportionate. Polymers such as polyethylene glycol, from detergents like Triton X-100 or from low-quality plastics and some hand creams, ionise extremely well and produce characteristic ladders of peaks spaced by 44 daltons that can dominate a spectrum and suppress peptide signals. Keratins from skin and hair are ubiquitous and are routinely included in search databases as contaminants so that they are recognised rather than misassigned. Residual SDS, even at trace levels, degrades chromatography. The trypsin itself autolyses and contributes peptides. None of this is exotic, and all of it is preventable with disciplined practice, but a single contaminated sample in a multiplexed set can compromise every sample it is pooled with. The special cases: plasma and low input Two kinds of sample push these principles to their limits and are worth treating separately because they shape so many current studies. Blood plasma and serum are the most clinically important and the most analytically hostile specimens in proteomics. Albumin alone makes up roughly half of the protein mass, and together with a small number of other high-abundance proteins such as immunoglobulins, transferrin, haptoglobin, fibrinogen and apolipoproteins, it accounts for most of the total. A shotgun analysis of neat plasma spends most of its instrument time on these proteins and typically quantifies several hundred proteins, not thousands. Three broad strategies extend depth. Immunoaffinity depletion columns remove the most abundant proteins before digestion, at the cost of expense, throughput, and the risk of co-depleting proteins bound to albumin or antibodies. Enrichment strategies, including perchloric acid precipitation and the use of engineered nanoparticles whose surfaces adsorb a protein corona that is enriched for lower-abundance species, compress the dynamic range in a different way; the nanoparticle approach, described by John Blume and colleagues in Nature Communications in 2020, has since been commercialised and used in large cohort studies. Finally, extensive fractionation, discussed in the next chapter, trades throughput for depth. Each strategy changes which proteins are measured and how reproducibly, so none of them can be treated as a transparent preprocessing step. A biomarker discovered after one kind of enrichment must be validated in a form that does not depend on it. Pre-analytical variation matters as much as analytical method in plasma work. The time a blood sample sits before centrifugation, the anticoagulant used, haemolysis, platelet contamination and freeze-thaw history all change the measured proteome. Philipp Geyer, Matthias Mann and colleagues catalogued proteins that act as markers of such problems, for example erythrocyte proteins as indicators of haemolysis, and proposed using them to flag compromised samples. In a clinical cohort where cases and controls were collected at different sites or at different times, these effects can generate differences that look like disease biology and are nothing of the kind. At the other extreme, the analysis of very small samples, from a few thousand cells down to single cells, turns sample loss into the dominant problem. Every surface a peptide touches adsorbs some of it, and at nanogram inputs the fraction lost to tubes, tips and columns can exceed the fraction measured. Low-input workflows therefore minimise transfers and volumes, performing lysis, digestion and sometimes labelling in the same nanolitre-scale well, and avoid clean-up steps where possible by using mass-spectrometry-compatible reagents. Single-cell proteomics, which has grown rapidly since the SCoPE-MS method was published by Bogdan Budnik, Nikolai Slavov and colleagues in 2018, relies on such miniaturised preparation. It is a separate field in its own right, but its lessons generalise: the less material there is, the more the preparation, rather than the instrument, sets the limit of detection. Designing the preparation, not just performing it Several design principles follow from this chapter, and they recur throughout the rest of the book. The first is that preparation defines the population being measured. A lysis buffer that fails to solubilise membranes produces a proteome depleted of membrane proteins, and no downstream method can recover them. When comparing studies, or when combining datasets, differences in preparation are often larger than differences in instrumentation. The second is that preparation is a source of batch effects. Samples prepared on different days, by different people, with different lots of trypsin or in different positions on a plate will differ systematically. The defence is randomisation and blocking: samples from each experimental group should be distributed across preparation batches so that batch and biology are not confounded, and the preparation order should be recorded so that batch can be modelled statistically later. A study in which all control samples were prepared in week one and all treated samples in week two cannot be rescued by any analysis. The third is that preparation choices interact with the acquisition strategy. Isobaric labelling requires that peptides be dissolved in an amine-free buffer, such as HEPES or triethylammonium bicarbonate, at a pH around eight to eight and a half, because TMT reagents react with primary amines; Tris buffer, which carries its own primary amine, will consume the reagent. Labelling efficiency must be checked, typically by a small test run searched with the tag set as a variable rather than fixed modification, because incompletely labelled peptides are invisible to the quantification. Label-free and DIA workflows avoid those constraints but depend even more heavily on run-to-run consistency of the digest, because every sample is measured separately and any difference in preparation propagates straight into the quantities. The fourth is that quality control starts here, not at the instrument. A standard digest, such as a commercial human cell-line digest, run at intervals alongside the study samples shows whether the chromatography and the instrument are stable; a pooled sample made from a small aliquot of every study sample, carried through preparation in each batch, shows whether the preparation is stable. The pooled reference plays an additional role in isobaric labelling designs, as a bridge between multiplexed sets, and Chapter 4 returns to it. Sample preparation is unglamorous, and its methods papers are less cited than those describing new instruments. But the depth, precision and validity of everything that follows are bounded by what happens in the first few hours of the workflow. The next chapter takes the peptide mixture that preparation produces and asks how to spread it out in time so that the mass spectrometer can see as much of it as possible. Chapter 2: Separation as a Resource A tryptic digest of a human cell line contains hundreds of thousands of distinct peptide species, spread across a concentration range of five or six orders of magnitude. If the whole mixture were sprayed into a mass spectrometer at once, the instrument would see only a dense forest of overlapping peaks dominated by the most abundant peptides, and the rest would be suppressed in the electrospray or lost beneath the noise. Separation solves this by spreading the mixture out in time, so that at any given moment only a manageable subset of peptides enters the instrument. It is best thought of not as a preliminary step but as a resource with a budget, measured in peak capacity, in instrument hours and in sample consumed. How that budget is spent determines the depth of the experiment and, increasingly, its throughput. Online chromatography: the conveyor belt into the instrument Almost all bottom-up proteomics couples reversed-phase liquid chromatography directly to electrospray ionisation. Peptides are loaded onto a column packed with small silica particles bearing C18 alkyl chains, in a mostly aqueous acidic solvent where they stick to the hydrophobic surface. A gradient of increasing organic solvent, usually acetonitrile, then elutes them roughly in order of hydrophobicity. The column outlet leads to a fine emitter at high voltage, where the liquid is dispersed into charged droplets that evaporate to release gas-phase peptide ions. Two properties of this arrangement shape everything else. The first is that electrospray behaves, over the relevant range, as a concentration-sensitive process: the signal depends on the concentration of analyte in the droplets rather than on the total amount delivered per unit time. Reducing the flow rate and the column diameter concentrates each peptide into a smaller volume and produces smaller initial droplets that ionise more efficiently. This is the reason the field adopted nanoflow chromatography, with columns of around 75 micrometres internal diameter running at a few hundred nanolitres per minute, as its default for discovery work. The sensitivity gain over conventional analytical flow is large, and for limited samples it is decisive. The second property is that chromatographic resolving power is finite. The useful measure is peak capacity, roughly the number of peaks that can be separated side by side within the gradient. It increases with column length, with smaller particle sizes that require higher pressures, and with longer gradients, but with diminishing returns: doubling gradient time does not double the number of peptides identified, because the gain in separation is partly offset by broader peaks and diluted signals. A typical nanoflow setup running a two-hour gradient on a 25 to 50 centimetre column may have a peak capacity of several hundred. Against a mixture of hundreds of thousands of peptides, that means many peptides still co-elute, and the mass spectrometer must sort them out by mass, by ion mobility and by fragmentation. Nanoflow's sensitivity comes at the cost of robustness. Nanolitre flow rates are hard to keep stable, the tiny emitters clog and degrade, retention times drift between runs and columns, and a single failure can interrupt a large sample series. For studies of hundreds or thousands of samples, where the limiting factor is not sample amount but reproducibility across weeks of continuous operation, several groups have argued for moving to higher flow. Yangyang Bian, Bernhard Kuster and colleagues, writing in Nature Communications in 2020, showed that microflow chromatography at several microlitres per minute, loading correspondingly more peptide, could deliver robust, reproducible analysis of thousands of proteomes with an acceptable loss of sensitivity. Clinical plasma proteomics projects in particular have moved toward higher flow rates and short gradients because they prize reliability over maximum depth per run. Throughput has become a design variable in its own right. A decade ago a typical deep proteome run lasted two to four hours. Dedicated chromatography systems now offer standardised methods defined by samples per day rather than gradient length. The Evosep One, described by Nicolai Bache, Philipp Geyer and colleagues in 2018, preforms the gradient at low pressure and stores it in a loop so that the high-pressure part of the run is short, which reduces the dead time between injections to a few minutes and allows standardised methods from tens to hundreds of samples per day. When paired with the fastest current instruments, short gradients of fifteen to thirty minutes now yield proteome depths that required hours of instrument time a few years ago. The Copenhagen group of Jesper Olsen, reporting on the Orbitrap Astral instrument in Nature Biotechnology in 2024, quantified about ten thousand human protein groups in half-hour runs and about seven thousand in five-minute runs using narrow-window DIA. Such figures are not typical of every laboratory or sample type, but they show how quickly the balance between depth and throughput has moved. Chromatography also introduces a subtle problem for quantification across runs. The retention time of a peptide shifts slightly between injections, and more between columns or laboratories. Label-free and DIA workflows depend on recognising the same peptide in every run, so they must align retention times. The indexed retention time, or iRT, concept introduced by Claudia Escher, Lukas Reiter and colleagues in 2012 expresses retention on a normalised scale anchored by a set of synthetic reference peptides spiked into every sample, which allows retention times from spectral libraries to be transferred to new runs. Modern software often uses endogenous peptides for the same purpose, and deep learning models now predict retention times directly from sequence with useful accuracy. Separation also governs a quieter effect that is easy to overlook: competition for charge in the electrospray. Peptides that elute together compete for the limited charge available on the surface of each droplet, and the more readily ionised species win. A peptide that ionises well in one sample may be partly suppressed in another if a highly abundant co-eluting peptide happens to be present there, for instance because the treatment induced a stress protein. Better separation reduces this competition by reducing the number of co-eluting species, which is one reason why the same peptide can give a more reliable quantity in a well-separated run than in a crowded one. It is also a reason why absolute signal intensities are not comparable between different peptides: each peptide has its own ionisation efficiency, determined by its sequence, and two peptides from the same protein, present at exactly the same molar amount, can differ in signal by more than an order of magnitude. Quantitative proteomics therefore almost always compares the same peptide across samples, rather than different peptides within a sample, and the methods for estimating absolute protein amounts from label-free data, such as summing intensities across all peptides of a protein or using only the most intense peptides, rely on averaging over this variation rather than removing it. Offline fractionation: buying depth with time When a single chromatographic dimension is not enough, the mixture can be split into fractions before online analysis, each of which is then run separately. The logic is simple. If the mixture is divided into twenty fractions that contain largely different peptides, each online run faces a far less complex mixture, and the instrument can sample deeper into each. The cost is equally simple: twenty times the instrument time per sample, and more handling. For fractionation to help, the offline separation must be orthogonal to the online one, meaning that it separates peptides by a property different from, or at least not strongly correlated with, hydrophobicity at low pH. The first widely used two-dimensional method was strong cation exchange, which separates peptides by charge. John Yates's group combined it with reversed-phase in a single capillary in the multidimensional protein identification technology, MudPIT, published by Michael Washburn, Dirk Wolters and John Yates in Nature Biotechnology in 2001, which identified well over a thousand yeast proteins at a time when that was remarkable. The method that displaced strong cation exchange for most deep proteome work is high-pH reversed-phase fractionation. It seems at first as though it should not work, since both dimensions are reversed-phase. The trick is that the selectivity of reversed-phase chromatography depends on the charge state of peptides, which changes with pH. At pH ten, acidic residues are deprotonated and basic ones largely neutral, so the elution order differs substantially from that at pH two or three. The orthogonality is imperfect, but the high-pH dimension has much higher resolving power than ion exchange and uses volatile, salt-free buffers that need no desalting. The remaining correlation between the dimensions is handled by concatenation, an approach described by Yuexi Wang and colleagues in 2011: the offline separation is collected into many small fractions, say ninety-six, and these are then pooled in a pattern in which early, middle and late fractions are combined, for instance fractions one, thirteen, twenty-five and so on into the first pooled fraction. Each pooled fraction then contains peptides spanning the whole hydrophobicity range of the second dimension, which uses the online gradient efficiently rather than concentrating peptides in one part of it. High-pH fractionation is how the deepest proteomes are made. Dorte Bekker-Jensen, Jesper Olsen and colleagues, in a 2017 Cell Systems paper, combined high-pH reversed-phase fractionation into dozens of fractions with fast data-dependent acquisition and reported identification of well over ten thousand protein-coding genes from human cell lines, approaching the complete expressed proteome. Such projects are resource maps rather than routine experiments, but they establish what is present to be measured and, importantly, they generate spectral libraries that make DIA analyses of single-shot runs more effective. Fractionation interacts strongly with the choice of quantification strategy, and this is one of the central design points of the book. In an isobaric labelling experiment, all samples in a multiplexed set are pooled before fractionation. One fractionated series of, say, twenty-four runs therefore delivers deep quantification for sixteen or eighteen samples at once. The instrument cost of fractionation is shared across the plex, which is why deep TMT experiments routinely fractionate and why they achieve depths that single-shot label-free analysis struggles to match. In a label-free or DIA experiment, each sample is measured separately, so fractionating each sample into twenty-four fractions multiplies instrument time for the whole study by twenty-four. For large cohorts that is prohibitive, and DIA studies therefore usually run each sample as a single shot and put their fractionation effort into building a library from a pooled sample, or into gas-phase fractionation, discussed below. A further design consideration is fraction-to-fraction variability. Because each fractionated sample is split into many runs, the quantity of a peptide measured in one sample depends on how it distributed across fractions, and small differences in fractionation between samples can introduce noise. In isobaric experiments this problem vanishes within a plex, because all samples experience the same fractionation physically together. In label-free designs with per-sample fractionation it must be controlled by summing signals across fractions and aligning carefully, which is one more reason large label-free studies avoid it. Separating in the gas phase Chromatography separates peptides in solution. Two further separation dimensions can act after ionisation, inside or just in front of the mass spectrometer, and both have become important in the last decade. Ion mobility separates ions by how quickly they drift through a gas under an electric field, which depends on their size, shape and charge. Two implementations dominate proteomics. High-field asymmetric waveform ion mobility spectrometry, or FAIMS, filters ions between the source and the mass spectrometer according to how their mobility changes between high and low fields; stepping through a small number of compensation voltages during a run acts as a form of online fractionation that removes singly charged background ions and reduces co-isolation. Alexander Hebert, Joshua Coon and colleagues showed in 2018 that FAIMS improved identification depth in single-shot analyses, and it has since been shown to reduce interference in isobaric labelling experiments as well. Trapped ion mobility spectrometry, or TIMS, as implemented on Bruker's timsTOF instruments, accumulates ions in a device where a gas flow is balanced against an electric field, then releases them in order of their mobility over tens of milliseconds. Coupled to a fast time-of-flight analyser, this allows parallel accumulation and serial fragmentation, or PASEF, in which the quadrupole jumps between precursors as they elute from the mobility device so that many precursors can be fragmented in a single mobility scan. Florian Meier, Matthias Mann and colleagues described PASEF in 2015 and 2018, and its extension to DIA, diaPASEF, in 2020. Ion mobility adds a coordinate to every peptide signal, and that extra coordinate is especially valuable for untangling the multiplexed spectra of DIA. Gas-phase fractionation is simpler and older. Rather than splitting the sample chemically, the same sample is injected several times and the instrument is told to examine a different slice of the precursor mass range each time. In DIA workflows this is an efficient way to build a deep, project-specific library: Brian Searle, Michael MacCoss and colleagues described in 2018 how a series of injections of a pooled sample, each covering a narrow mass range with very narrow isolation windows, could generate a chromatogram library that is directly matched to the experimental runs because it was collected on the same instrument under the same chromatography. Budgeting separation Several practical principles follow. Separation effort should be spent where the question needs it. A study seeking to catalogue low-abundance transcription factors in a cell line needs depth, which calls for fractionation, long gradients or both. A study of three hundred patient plasma samples seeking robust associations with a clinical variable needs reproducibility and throughput, which calls for short, robust single-shot runs, perhaps with higher flow. A time course with a modest number of samples and a need for complete data across all of them may be well suited to a single multiplexed set with extensive fractionation. The same instrument can serve all three, but not with the same method. Instrument time is a finite, expensive resource, and it should be counted at the design stage. A cohort of five hundred samples run in single-shot DIA at sixty samples per day consumes a little over a week of continuous acquisition, plus quality controls and blanks. The same cohort run with twenty-four fractions per sample on a two-hour gradient would need years. A deep multiplexed design would need roughly thirty sets of sixteen samples, each with its own fractionated series. Comparing these budgets at the planning stage, before any samples are prepared, prevents the common failure in which a study design is fixed first and discovered later to be unaffordable. Run order matters. Chromatographic performance drifts over hundreds of injections as columns age and emitters degrade, and carryover of abundant peptides from one run to the next is real. Samples should be injected in randomised or balanced order so that drift does not align with experimental groups, blanks should be interleaved where carryover is a concern, and quality-control standards should be run at regular intervals so that drift can be detected and, if necessary, corrected. Finally, separation cannot be optimised in isolation from acquisition. A short gradient produces narrow chromatographic peaks, perhaps a few seconds wide, which demands a fast instrument to sample each peak several times. A long gradient spreads peptides out and relaxes that constraint but reduces throughput. The next chapter turns to the mass spectrometer itself, and to the question of how an instrument that can examine only a few precursors at a time decides which ones to look at. Chapter 3: Inside the Mass Spectrometer A mass spectrometer does two things that matter for proteomics. It measures the mass-to-charge ratio of ions with great accuracy, and it can isolate chosen ions, break them apart and measure the pieces. The first operation, recorded as an MS1 or survey spectrum, says which peptide ions are present at a given moment and how intense they are. The second, recorded as an MS2 or tandem mass spectrum, provides the fragment pattern from which a peptide's sequence can be inferred. Every acquisition strategy, whether data-dependent, isobaric or data-independent, is a policy for deciding when to do which of these operations and on which ions. To understand why TMT and DIA behave as they do, it helps first to understand the machinery and the older policy they both set out to improve. Analysers, speed and fragmentation The core of a mass spectrometer is its mass analyser, the device that sorts ions by mass-to-charge ratio. Modern proteomics instruments are hybrids that combine two or three analysers, each doing what it does best. Table 1 summarises the main types in use for bottom-up work. Table 1. Mass analysers commonly used in bottom-up proteomics and their characteristic roles. Analyser Operating principle Main strength Typical role Quadrupole Oscillating electric fields pass a narrow m/z range Fast, selective isolation Precursor selection before fragmentation Linear ion trap Ions stored and ejected by m/z Sensitive, fast, flexible Fragment scans, MS3 selection Orbitrap Ions orbit a spindle electrode; frequency gives m/z High resolution and mass accuracy Survey scans, reporter ion and fragment scans Time-of-flight Flight time over a fixed path gives m/z Very fast acquisition Fragment scans, especially with ion mobility Astral Lossless multi-reflection time-of-flight path Speed with high sensitivity and resolution Fast MS2 for narrow-window DIA and DDA Three performance characteristics matter most. Resolution is the ability to distinguish ions of nearly identical mass-to-charge ratio. It determines whether two co-eluting peptides with similar masses appear as separate peaks and, critically for isobaric labelling, whether reporter ions that differ by a few thousandths of a dalton can be told apart. Mass accuracy is how close the measured mass is to the true mass; modern Orbitrap and time-of-flight instruments routinely achieve errors of a few parts per million, which sharply limits the number of candidate peptides that could explain a given precursor. Speed, expressed as the number of MS2 spectra that can be acquired per second, sets how many peptides can be sampled during the few seconds each one takes to elute. These properties trade against one another. In an Orbitrap, higher resolution requires a longer transient, the period over which the image current of the orbiting ions is recorded, and therefore a slower scan. Sensitivity also depends on how long ions are accumulated before each scan: collecting ions for longer gives better spectra of low-abundance peptides but slows the cycle. Instrument generations since the Orbitrap was introduced commercially in 2005 have progressively relaxed these trade-offs, and the introduction of the Orbitrap Astral in 2023, which pairs an Orbitrap for survey scans with the fast, sensitive Astral analyser for fragment scans at rates reported to exceed two hundred spectra per second, changed the practical ceiling on how many peptides can be fragmented in a run. Bruker's timsTOF line, pairing trapped ion mobility with time-of-flight, achieved high fragmentation rates by a different route. The competition between these platforms continues, and any statement about which instrument is best has a short shelf life. What endures is the logic: an acquisition strategy is a way to spend a fixed budget of ions and time. Fragmentation turns a precursor mass, which is compatible with many possible sequences, into a pattern that points to one. The dominant method is collision-induced dissociation, in which peptide ions are accelerated into an inert gas such as nitrogen or argon; the collisions deposit energy that breaks the peptide backbone, predominantly at the amide bonds. In the higher-energy collisional dissociation implemented on Orbitrap instruments, known as HCD, fragmentation takes place in a dedicated collision cell and all fragments are measured, including the low-mass region where TMT reporter ions appear. Earlier ion trap collision-induced dissociation had a low-mass cut-off that made reporter ions hard to observe, which is one reason HCD became the standard for isobaric labelling. When a peptide breaks at an amide bond, the charge can stay with either piece. Fragments that contain the amino terminus are called b ions, those that contain the carboxyl terminus y ions. A complete ladder of b or y ions spells out the sequence, since successive fragments differ by the mass of one amino acid residue. Real spectra are incomplete and noisy: some bonds rarely break, especially next to proline, some fragments carry multiple charges, neutral losses of water or ammonia add peaks, and the intensities of fragments vary in ways that depend on sequence. Tryptic peptides, with their basic carboxyl terminus, tend to produce strong y-ion series, which is one reason they are so easy to identify. Alternative fragmentation methods exist for specific problems. Electron transfer dissociation, developed by Donald Hunt, Joshua Coon, John Syka and colleagues in 2004, transfers an electron to a multiply charged peptide and breaks a different bond along the backbone, producing c and z ions. It tends to preserve labile modifications such as phosphorylation and glycosylation that collisional methods strip off, and it works well for highly charged, long peptides. Hybrid methods combine electron transfer with supplemental collisional activation. For deep quantitative profiling of unmodified peptides, however, HCD is the default, and the rest of this book assumes it unless otherwise stated. Data-dependent acquisition and its discontents For most of the history of shotgun proteomics, the policy for deciding what to fragment was data-dependent acquisition. The instrument records a survey scan, detects the peptide-like features in it, and selects the most intense, typically the top ten to twenty or as many as it can fit into a fixed cycle time. For each, the quadrupole isolates a narrow window of mass-to-charge ratio around the precursor, usually one to two units wide, the ions in that window are fragmented, and the MS2 spectrum is recorded. The cycle then repeats. Dynamic exclusion prevents the same precursor from being picked again for a set period, typically tens of seconds, so that the instrument moves on to less abundant species rather than repeatedly fragmenting the same abundant ones. Filters reject singly charged ions, which are mostly contaminants rather than tryptic peptides. The strengths of the approach are real. Each MS2 spectrum is, in principle, derived from a single precursor whose mass is known, which makes it easy to interpret. Search engines designed in the 1990s were built for exactly this kind of data, and the database search methods described in Chapter 6 remain the backbone of identification. The weaknesses emerged as the field tried to scale. The first is undersampling. A single shotgun run contains far more peptide species than the instrument can fragment. Annette Michalski, Jürgen Cox and Matthias Mann estimated in a 2011 Journal of Proteome Research paper that more than one hundred thousand peptide features were detectable in the survey scans of a single run of a complex human digest, but that the majority were inaccessible to data-dependent fragmentation, either because the instrument did not have time to pick them or because they were too weak to yield identifiable spectra. The run-to-run consequence is stochastic sampling. Which peptides reach the top of the intensity list at the moment the instrument is looking depends on small variations in chromatography and ion accumulation. A moderately abundant peptide may be fragmented in one run and not in the next. Across a study of many samples this produces a data matrix riddled with missing values, not because the peptide is absent from the sample but because the instrument happened not to select it. The second weakness is co-isolation. The isolation window around the selected precursor is narrow, but at any moment in a complex run it usually contains other peptide ions as well, some of them near the noise level of the survey scan. Their fragments contaminate the MS2 spectrum. Such chimeric spectra reduce identification scores and, when the quantity is read from the fragment spectrum, as in isobaric labelling, contaminate the measurement. This problem is central to Chapter 4. A short calculation makes the constraints concrete. Suppose a peptide elutes from the column over a chromatographic peak about fifteen seconds wide at its base, which is typical of a one- to two-hour nanoflow gradient; on a short high-throughput gradient the peak may be only five seconds wide. At any moment during a busy part of the gradient, perhaps several hundred peptide features are visible in the survey scan. If the instrument can acquire twenty fragment spectra per second, and each cycle consists of one survey scan followed by as many fragment scans as fit into a second, then in the fifteen seconds that one peptide is eluting the instrument can fragment around three hundred precursors, most of them only once because of dynamic exclusion. That sounds generous until one recalls that the same fifteen seconds contain several hundred to a thousand co-eluting features, many of them too weak to yield a useful spectrum with a short ion accumulation time. Doubling the speed of the instrument helps, but if the injection time must be shortened to achieve that speed, weak precursors produce even poorer spectra. The same arithmetic applies differently to quantification. To integrate a chromatographic peak reliably from MS1 survey scans, the peak must be sampled several times; a common rule of thumb in targeted and DIA work is at least six to ten points across the peak. With a fifteen-second peak, that requires a survey scan every two seconds or faster, which is easy. With a five-second peak it requires a cycle time under a second, which constrains how much fragmentation can be fitted between surveys. The move toward short gradients for throughput therefore pushes instruments toward higher speed, and it is no accident that the fastest current analysers were developed at the same time as high-throughput chromatography. Ion accumulation introduces a second constraint. Orbitrap instruments collect ions in a trap before each scan, controlled by a target number of charges called the automatic gain control setting and a maximum injection time. The target exists because too many ions in the analyser distort the measurement through space-charge effects, degrading mass accuracy. The consequence is that each spectrum has a finite dynamic range: if a very abundant ion fills most of the available capacity, weaker ions in the same scan are represented by few charges and measured imprecisely. This is one reason why dynamic range within a single survey scan is several orders of magnitude narrower than the dynamic range of the proteome, and why separation, which reduces the number of ions competing for each scan, matters so much. Label-free quantification in data-dependent mode Even without labels, data-dependent runs can be used to quantify. The crudest method, spectral counting, simply counts how many MS2 spectra were matched to each protein, on the reasoning that abundant proteins yield more peptides and more repeated selections. It is easy and surprisingly informative for large differences, but it saturates for abundant proteins, is noisy for scarce ones and is distorted by dynamic exclusion. The more precise approach uses the MS1 signal: once a peptide is identified from its MS2 spectrum, the area under its chromatographic peak in the survey scans is integrated, and that area serves as its quantity. Peptide quantities are then combined into protein quantities. To address missing values, software can transfer identifications between runs. The widely used MaxQuant package implemented this as match between runs: if a peptide was identified by MS2 in one run, the software looks in other runs for an MS1 feature with the same accurate mass and a retention time that, after alignment, is close enough, and assigns the identity even though no MS2 spectrum was taken. The MaxLFQ algorithm described by Jürgen Cox and colleagues in 2014 then computes protein intensities from pairwise peptide ratios in a way designed to be robust to missing values and to fractionation. Match between runs substantially reduces missing values, but it introduces a new kind of error. A transferred identification rests on mass and retention time alone, and in a crowded run there may be several features that fit. Controlling the error rate of these transfers has been a recurring concern, and several groups have proposed methods to estimate it, for instance by transferring decoy identifications alongside real ones and counting how often they are matched. What the discontents set in motion The two strategies that are the core of this book can be understood as different answers to the problems of data-dependent acquisition. Isobaric labelling accepts the data-dependent paradigm, including its stochastic sampling, but changes what a single selection achieves. By pooling many samples before analysis, it ensures that whenever a peptide is picked, it is quantified in all of them simultaneously. The missing-value problem is solved within a multiplexed set, though it reappears between sets. Co-isolation, however, becomes more damaging, because the quantity itself now comes from the fragment spectrum. Data-independent acquisition abandons the idea of choosing precursors. By fragmenting everything in systematic windows, it guarantees that every detectable peptide is sampled in every run, removing stochastic sampling at the acquisition stage. The price is that co-isolation is no longer an occasional contamination but the normal state of every spectrum, and the burden of sorting fragments to their precursors shifts entirely to software. Both answers are clever, both have limitations, and the next two chapters examine each in turn. The distinction between them is sometimes presented as labelled versus label-free, but that obscures more than it reveals. The deeper distinction is between reducing the number of runs and making every run complete. Hybrid methods that combine labelling with data-independent acquisition, discussed in Chapter 5, show that the two ideas are not mutually exclusive. One final point about instruments deserves emphasis, because it affects how published comparisons should be read. Each acquisition strategy was developed on, and optimised for, particular instrument architectures, and its relative performance changes with each hardware generation. Isobaric labelling with synchronous precursor selection and MS3 was developed on Orbitrap tribrid instruments that have a linear ion trap; diaPASEF was developed on timsTOF instruments; narrow-window DIA was enabled by the Astral analyser. A comparison between methods performed on one instrument generation is informative about the principles but not always about the present. Readers evaluating the literature, or planning a study, should ask which instrument was used, in what year, and whether the conclusion depends on a speed or sensitivity limit that has since moved. Hashtags: #ProteomicsByDesign #QuantitativeProteomics #BottomUpProteomics #TandemMassTags #TMTProteomics #IsobaricLabeling #DataIndependentAcquisition #DIAProteomics #DataDependentAcquisition #MassSpectrometry #ShotgunProteomics #RatioCompression #SPSMS3 #SpectralLibraries #LibraryFreeDIA #IonMobility #DiaPASEF #ProteinQuantification #PeptideIdentification #TargetDecoyStrategy #FalseDiscoveryRate #QuantitativeBioinformatics #DifferentialAbundance #ProteomicsExperimentalDesign #FutureOfQuantitativeProteomics

  • Proteogenomics (Unifying Genomics, Transcriptomics, and Mass Spectrometry)

    Download the Book (PDF): Introduction A mass spectrometer does not read proteins. It weighs fragments of them. When a peptide is broken apart inside the instrument, the resulting spectrum is a list of masses and intensities, and nothing in that list announces which gene the peptide came from. To turn a spectrum into a biological statement, a computer compares it with the spectra that would be expected from a set of candidate sequences and reports the best match. The candidates come from a file, a protein sequence database, which someone had to decide to build in a particular way. That file is the quiet centre of this book. For most of the history of shotgun proteomics, it was taken for granted. A laboratory downloaded the reference proteome for its organism, searched its spectra against it, and reported the proteins it found. This works well for a great deal of biology, and it will continue to. But it builds in an assumption that is easy to forget: that every protein worth finding is already written down. A peptide that is not in the database cannot be identified, however clean its spectrum. It will either go unmatched or, worse, be assigned to the closest sequence that is in the file, and reported as something it is not. Proteogenomics is the discipline that refuses this assumption. It uses genomic and transcriptomic data, usually from the same sample the proteins came from, to write a better database: one that contains the patient's own single-nucleotide variants, the splice junctions actually used in that tissue, the transcripts assembled from its RNA, the open reading frames that ribosomes were observed to translate. Mass spectrometry then serves as the arbiter. If a peptide carrying a predicted amino-acid change is detected, the variant is not merely present in DNA; it is expressed as protein. If a peptide spans an exon junction never seen in the reference annotation, a new isoform is supported at the protein level. If a peptide maps to a region annotated as non-coding, the annotation is challenged. The term covers two traffic directions. In one, protein data correct and extend genome annotation: peptides are used to confirm gene models, find missed genes and fix start sites. This was the original meaning, and it remains central to annotating newly sequenced organisms. In the other, genomic data make proteomics more complete and more personal: sample-specific sequences are added so that the proteins peculiar to one tumour, one cell line or one person can be seen at all. The cancer-focused work of the past decade, including the large programmes that profiled hundreds of tumours with both sequencing and deep proteomics, belongs mainly to this second direction. In practice the two meet in the same computational pipeline. The argument of this book The controlling idea here is simple to state and easy to underestimate. In proteogenomics, the sequence database is the hypothesis, and every discovery is only as trustworthy as the design of that database and the validation that follows. Adding sequences buys the possibility of discovery. It also costs statistical power, invites false matches and opens the door to artefacts that masquerade as biology. The whole craft lies in managing that trade: expanding the search space far enough to see what the reference misses, and no further than the evidence can support, then testing every surprising match as if it were wrong until shown otherwise. This framing matters because the field has repeatedly been burned by forgetting it. Early six-frame translation searches reported novel peptides that later proved to be modified or mis-assigned canonical ones. Claims about how much of the immunopeptidome arises from exotic processing mechanisms were revised sharply once the search strategies behind them were scrutinised. Headline counts of "novel proteins" have tended to shrink when stricter evidence standards were applied. None of this means proteogenomics is unreliable. It means the reliability is manufactured, step by step, by choices that a careful analyst can make well or badly. What the book covers The chapters follow the logic of a real analysis. The first explains why reference proteomes are incomplete, and why that incompleteness is structured rather than random: it is concentrated in short proteins, non-canonical reading frames, individual variation and tissue-specific isoforms. The second explains how mass spectrometry actually identifies peptides, with particular attention to the target-decoy method and the false discovery rate, because everything later depends on understanding what a 1 percent threshold does and does not promise. The third chapter turns to construction: how variant calls, RNA sequencing assemblies, splice-junction graphs and ribosome profiling are converted into protein sequences, and what each source contributes and costs. The fourth faces the statistical consequence of larger databases, explaining why a novel peptide passing a global threshold can still be much more likely to be wrong than a canonical one, and what the accepted remedies are. Four chapters then follow the applications. One concerns alternative splicing and the real, somewhat sobering evidence about how many alternative isoforms are detectable as protein. One concerns non-canonical proteins: the short open reading frames, upstream reading frames and translated "non-coding" RNAs that ribosome profiling has revealed, and the difficulty of confirming them by mass spectrometry. One reviews what the large tumour cohort studies, which measured genome, transcriptome, proteome and phosphoproteome in the same specimens, actually taught about how genomic change reaches the protein level. One concerns personalised tumour antigens, where the stakes are clinical, the peptides are displayed on human leukocyte antigen molecules rather than produced by trypsin, and the question is not only whether a peptide exists but whether a T cell can see it. The final chapter gathers the validation protocols that apply across all three: peptide-centric re-searching against every competing explanation, spectral comparison with synthetic standards and predicted spectra, retention-time consistency, and a tiered way of stating how confident a claim deserves to be. Who this is for The reader imagined here is scientifically literate but not necessarily a specialist in both halves of the field. Many readers arrive from one side: a genomicist who has variant calls and wants to know whether they are expressed; a proteomicist who has unexplained spectra and wants to know what they might be; a cancer immunologist who wants to know which neoantigen predictions to trust; a bioinformatician asked to build the pipeline. The book tries to give each enough of the other side to reason well, without pretending to be a software manual. Tools are named where they illustrate a principle, and the principles are what should outlast the tools. No knowledge of statistics beyond the idea of a threshold is assumed, but the book does not avoid the statistics, because the statistics are where most proteogenomic claims succeed or fail. Nor does it avoid the chemistry of fragmentation and the arithmetic of residue masses, because a surprising number of apparent discoveries dissolve once someone checks whether two amino acids, or an amino acid and a modification, weigh the same. A note on evidence Proteogenomics sits at the join between two data types that are each noisy in their own ways. Sequencing errors, alignment artefacts and mis-called variants propagate into the database; spectral noise, co-isolation and unexpected modifications propagate into the matches. When the two are combined, errors do not cancel. They compound, because a false sequence in the database gives a noisy spectrum something to be falsely matched to. The recurring advice in these pages is therefore conservative: prefer sample-specific evidence to generic expansion, report novel findings under separate and stricter thresholds, and treat every novel peptide as a claim requiring a second, independent line of support. This is not caution for its own sake. The payoff of the discipline, when it is practised well, is considerable. It has shown which cancer mutations reach the protein level and which do not, revealed signalling consequences of somatic alterations that genomics alone could not predict, supported the translation of reading frames that annotation had missed, and identified tumour-specific peptides actually displayed to the immune system. Those results are credible because the people who produced them built their databases carefully and tested their matches hard. The chapters that follow set out how. Chapter 1: The Incomplete Reference Every protein sequence database is a model of what a cell can make. The reference proteomes distributed by UniProt, Ensembl and RefSeq are extraordinarily good models, built over decades from cDNA cloning, comparative genomics, manual curation and, increasingly, protein-level evidence. They are also models of a particular kind: a consensus drawn from many individuals, weighted toward long, conserved, well-studied proteins, and organised around rules for what counts as a gene. Understanding where those rules leave gaps is the first step in deciding what a customised database should add. How genes get annotated Genome annotation is an inference from evidence. Annotators align transcripts and known proteins to the genome, look for open reading frames with plausible start and stop codons, use conservation across species to judge whether a reading frame is under selection for its protein product, and then decide. The GENCODE project, which provides the reference gene set used by Ensembl for human and mouse, combines automated pipelines with manual curation by the HAVANA group, and its releases describe tens of thousands of protein-coding and non-coding genes and a much larger number of transcripts. Two design choices in this process shape proteogenomics directly. First, annotation historically required a minimum open reading frame length, commonly around 100 codons, to call a protein-coding gene without additional evidence. The reason was statistical: random sequence produces short open reading frames by chance all the time, so length was a cheap filter against false genes. The cost was that genuine short proteins were systematically excluded. Second, annotation relied heavily on evolutionary conservation. A reading frame whose protein-coding character was conserved across mammals was easy to accept; a young, lineage-specific or poorly conserved one was not. Both filters are sensible, and both create blind spots in exactly the places proteogenomics now looks. Protein-level evidence has itself reshaped the gene count. Ezkurdia and colleagues, combining large proteomics datasets with conservation and other evidence, argued in 2014 that there may be as few as 19,000 human protein-coding genes, fewer than annotations then listed, because many annotated coding genes showed no peptide evidence and features suggesting they were not truly coding. The point is not the precise number but the direction: annotation is revised in both directions, removing genes that fail to earn their status as well as adding new ones, and mass spectrometry is one of the tools that does the revising. What mass spectrometry has already seen In 2014 two groups published draft maps of the human proteome in the same issue of Nature. Kim and colleagues profiled a panel of adult and fetal tissues and primary cells; Wilhelm and colleagues assembled a large compendium of their own and public data into the ProteomicsDB resource. Together they showed that deep, tissue-diverse shotgun proteomics could detect peptides from the large majority of annotated protein-coding genes. They also reported peptides that did not match the reference: some from putative non-coding regions, some from alternative reading frames, some from unannotated exons. Several of those claims were later questioned, which is itself instructive, but the broad lesson held. A deep proteome contains signal that the reference database cannot explain. The Human Proteome Project, coordinated by the Human Proteome Organization, has tracked this systematically through the neXtProt knowledgebase, grading each human protein by its level of existence evidence. By 2020, Adhikari and colleagues reported that credible mass spectrometry or other protein-level evidence had been obtained for more than 90 percent of the predicted human proteins. The remaining "missing proteins" are not randomly distributed: they include olfactory receptors and other proteins expressed in few cells, membrane proteins with few tryptic sites, and proteins expressed only in specific developmental stages or tissues. The project's guidelines for claiming detection, discussed in the final chapter, are some of the most useful statements anywhere of what evidence a novel protein claim requires. Four kinds of gap For the purposes of database design it helps to sort the gaps in the reference into four kinds, because each is filled from a different data source and each carries a different risk. Individual variation. The reference proteome is essentially one sequence per protein, reflecting a reference genome assembled from a small number of donors. Any real individual differs from it at millions of positions, a few tens of thousands of which fall in coding sequence and change an amino acid. Population resources such as gnomAD catalogue these variants across many thousands of people. A peptide covering a missense variant will not match the reference sequence exactly, so in a standard search it goes unidentified or is matched incorrectly. For most applications the loss is small, since other peptides from the same protein are still found. For some it is the whole point: the protein product of a germline disease allele, or a somatic mutation in a tumour, is precisely the peptide a standard search cannot see. Somatic alteration. Cancer genomes add their own changes: point mutations, small insertions and deletions that shift the reading frame, gene fusions that join two coding sequences, and large-scale rearrangements. Frameshift insertions and deletions produce stretches of entirely novel amino-acid sequence downstream of the event, often until a new stop codon is reached. Fusions produce junction peptides that exist in no normal cell. These are the raw material of tumour-specific antigens, and none of them appear in any reference. Isoform diversity. Most human multi-exon genes produce more than one transcript through alternative splicing, alternative promoters and alternative polyadenylation. Reference databases include many annotated isoforms, but not every junction used in a given tissue, and certainly not the aberrant splicing that is common in tumours, where mutations in splicing factors such as SF3B1 generate recurrent novel junctions. A peptide that spans a novel exon-exon junction is invisible unless the junction is in the database. Non-canonical translation. Ribosomes translate more than annotated coding sequences. They initiate at upstream open reading frames in 5′ untranslated regions, at near-cognate start codons such as CUG, in reading frames overlapping known coding sequences, and within RNAs annotated as long non-coding. Many of the resulting products are short, some are unstable, and some have demonstrable functions. They are the least represented class in reference databases, partly because of the length and conservation filters described above. Why transcripts are not proteins A natural response to these gaps is to ask why proteomics is needed at all. If RNA sequencing shows the variant allele, the novel junction or the translated reading frame, why not accept it? The answer is that each step between DNA and functional protein is regulated and lossy. A somatic variant can be present in DNA but lie in an allele that is not expressed. A transcript can be abundant but subject to nonsense-mediated decay, or retained in the nucleus, or translated poorly. A reading frame can be engaged by ribosomes, as ribosome profiling shows, yet yield a product degraded almost as soon as it is made. Large studies that measured both mRNA and protein in the same tumours have repeatedly found that the correlation between them, gene by gene across samples, is positive but modest and highly variable. The first large proteogenomic study of colorectal cancer from the Clinical Proteomic Tumor Analysis Consortium, published by Zhang and colleagues in 2014, found that mRNA levels were limited predictors of protein abundance, and that many copy-number alterations visible in the genome had surprisingly weak effects at the protein level, with notable exceptions in a small number of amplified regions. Proteins are also the molecules drugs bind and immune receptors recognise. A T cell does not see a mutation; it sees a short peptide displayed by a human leukocyte antigen molecule on the cell surface. A kinase inhibitor does not act on a gene; it acts on a folded protein, whose activity may be set by phosphorylation that no sequencing assay measures. For questions like these, protein-level evidence is not a confirmation of the genomic result but a different and often more decisive observation. Why proteins cannot simply replace sequencing The complementary question is why proteomics needs genomics. The answer lies in coverage and sensitivity. A shotgun experiment samples peptides stochastically and favours abundant proteins; the dynamic range of protein concentrations in a cell or tissue spans many orders of magnitude, and in plasma considerably more. For any given protein only a subset of its peptides is observable, because some are too short or too long after digestion, some ionise poorly, and some carry modifications. The probability that the one peptide covering a specific variant, junction or novel reading frame is among those detected is often low. Mass spectrometry also cannot easily distinguish sequences that are identical in mass. Leucine and isoleucine have the same elemental composition; glutamine and lysine differ by only about 0.036 daltons; the dipeptide glycine-glycine has the same composition as asparagine. Without a candidate sequence to test, many spectra are consistent with several sequences. De novo sequencing, which reads peptide sequences directly from spectra without a database, has improved markedly with deep learning, but it still struggles with incomplete fragmentation and these ambiguities, and it is most powerful when combined with a database rather than replacing one. So each technology supplies what the other lacks. Sequencing proposes what could exist, with high sensitivity and near-complete coverage of the transcribed genome. Mass spectrometry tests which of those possibilities exist as protein, with lower sensitivity but direct physical observation of the molecule. The union is more informative than either, but only if the proposal step does not overwhelm the testing step, a tension that runs through every chapter that follows. Beyond the human genome It is worth remembering that proteogenomics began not with personalised medicine but with genome annotation of organisms for which no mature reference existed. When a bacterial, fungal, plant or parasite genome is first sequenced, its gene models are predicted computationally, and prediction errors are common: missed small genes, wrong start codons, genes predicted on the wrong strand, frameshifts introduced by sequencing errors in homopolymer runs. Searching the organism's own proteome against a translation of its whole genome allows peptides to vote on the gene models. A peptide found upstream of a predicted start codon indicates that translation begins earlier; a peptide mapping to an unannotated intergenic region suggests a missed gene; a set of peptides that straddle a predicted frameshift suggests the sequence, not the organism, is at fault. In compact prokaryotic genomes this strategy is powerful because the search space is small. A few million bases translated in six frames yields a database of manageable size, and most of the genome codes for protein anyway, so the expansion from the predicted proteome to the full translation is modest. The approach has been used to refine annotations for many bacteria and archaea, and it pairs naturally with methods that enrich protein amino termini, which pinpoint where translation actually starts. The same strategy applied to the human genome, which is several hundred times larger and overwhelmingly non-coding, behaves very differently, a contrast that will matter a great deal in the chapters on database construction and statistics. What counts as novel The word "novel" carries a lot of weight in proteogenomic papers, and it helps to be exact about it before going further. A peptide can be novel relative to the particular database used for a standard search, relative to all major public references, or relative to the scientific literature. These are different claims. A peptide absent from a curated canonical database may be present in an isoform database, in a variant database or in an older annotation release; a peptide absent from all current references may already have been reported in a ribosome profiling study or a previous proteogenomic analysis. A second distinction concerns what the novel peptide implies. Some novel peptides imply only a small change to a known protein: a single amino-acid substitution, a different exon junction, an extended amino terminus. Others imply an entirely new protein product: a reading frame in an untranslated region, in a long non-coding RNA or in an alternative frame of a known gene. Still others imply a molecular event that exists only in the sample: a somatic frameshift or a fusion junction. Each category carries a different prior probability of being real, and the book will argue repeatedly that this prior should shape how much evidence is demanded. A variant peptide predicted from a high-confidence DNA call and supported by RNA reads is a modest claim. A peptide implying a previously unknown protein from a non-coding region is an extraordinary one. Finally, novelty at the peptide level does not automatically transfer to the protein level. A single peptide shows that a particular sequence was present in the sample; it does not show which full-length protein it came from, whether that protein is stable, or whether it does anything. Much of the careful language in the proteogenomic literature, such as speaking of "peptide evidence for translation" rather than "a new protein", reflects this gap. It is a good habit to adopt from the start. The personal proteome One more gap deserves emphasis because it motivates the most ambitious applications. Every individual, and every tumour, has a proteome that differs from the reference in ways that matter to that individual. A pharmacogenomic variant may change an enzyme's active site; a germline variant may create or destroy a cleavage site that changes which peptides are produced; a tumour's somatic mutations may create epitopes that its immune system can target. A sample-specific database, built from that sample's own genome and transcriptome, represents this personal proteome directly. The idea is appealing and the benefits are real, but it also shifts the burden of proof. A reference database has been checked by many people over many years. A personal database was built last week by one pipeline from one sequencing run. Its errors are private too. Every variant call, every assembled transcript and every predicted junction is a candidate, not a fact, and the proteomic search is only one of several tests it must pass. Holding both halves of that picture in view, the promise of seeing the personal proteome and the fragility of the database that makes it visible, is the habit of mind proteogenomics requires. Chapter 2: How a Spectrum Becomes a Peptide Before deciding what to put in a database, it is necessary to understand exactly how the database is used. The details of peptide identification are often treated as a black box by people who use its output, and that is where most proteogenomic mistakes begin. This chapter walks through the path from protein sample to reported peptide, concentrating on the points where the database exerts its influence. From protein to spectrum In a typical bottom-up experiment, proteins extracted from cells or tissue are denatured, their cysteines reduced and alkylated, and the mixture digested with trypsin. Trypsin cuts after lysine and arginine, except, as a rule, when the next residue is proline. The result is a very complex mixture of peptides, most between about seven and thirty residues long, which is separated by reversed-phase liquid chromatography and sprayed into the mass spectrometer by electrospray ionisation. The instrument performs two kinds of measurement in alternation. A survey scan, called MS1, records the mass-to-charge ratios of the intact peptide ions eluting at that moment. The instrument then selects precursor ions, isolates them within a narrow window, fragments them, usually by collisions with gas in a process called higher-energy collisional dissociation, and records the masses of the fragments in an MS2 scan. Peptide bonds tend to break along the backbone, producing two series of ions: b ions, which contain the amino terminus, and y ions, which contain the carboxy terminus. The mass differences between successive ions in a series correspond to residue masses, so a complete series spells out the sequence. Real spectra are rarely complete. Some bonds break reluctantly, especially next to proline or in certain residue pairs. Neutral losses of water or ammonia add peaks. Other peptides co-isolated in the same window contribute fragments of their own. Modifications, whether biological such as phosphorylation or introduced during sample preparation such as oxidation of methionine, shift masses. A single experiment on a modern instrument records tens of thousands to hundreds of thousands of MS2 spectra, and a sizeable fraction of them are never confidently identified. In data-dependent acquisition, the mode described above, the instrument chooses which precursors to fragment on the fly, usually the most intense. In data-independent acquisition, it instead fragments everything within successive wide windows across the mass range, producing composite spectra that must be deconvolved computationally. Data-independent methods give more consistent quantification across samples and have become common in large cohort studies, but their identification strategies depend even more heavily on prior knowledge of what peptides to look for, typically in the form of spectral libraries or predicted spectra. Most proteogenomic discovery has so far been done with data-dependent acquisition, for reasons that will become clear. Database searching The dominant method for identifying peptides from MS2 spectra is sequence database searching. The search engine digests every protein in the database in silico according to the enzyme rules, computes the mass of every resulting peptide, and for each observed spectrum retrieves the candidate peptides whose masses lie within a tolerance of the measured precursor mass. For each candidate it predicts the fragment ions expected, compares them with the observed peaks and computes a score. The best-scoring candidate becomes the peptide-spectrum match, or PSM, for that spectrum. Several parameters control how many candidates each spectrum faces. The precursor mass tolerance on modern high-resolution instruments is typically a few parts per million, which is narrow; the number of missed cleavages allowed, usually two, multiplies the number of peptides per protein; and every variable modification allowed, such as methionine oxidation or amino-terminal acetylation, multiplies candidates further. Semi-specific or non-specific enzyme settings, needed when peptides do not arise from clean tryptic cleavage, expand the space by an order of magnitude or more. This matters because the score of the best match is only meaningful relative to the scores of the competitors it beat. The more candidates a spectrum is compared with, the higher the best random score it will encounter. Search engines differ in how they score. Early tools such as SEQUEST used cross-correlation between observed and predicted spectra; Mascot uses a probability-based score; X!Tandem, Comet, MS-GF+ and Andromeda in MaxQuant each have their own scoring functions and statistical models. MSFragger, introduced by Kong and colleagues in 2017, uses a fragment-ion index to make searches dramatically faster, which made practical the "open" searches that allow any mass shift on the precursor and thereby reveal unexpected modifications. That capability turns out to be important in proteogenomics, because many apparent novel peptides are really known peptides carrying a modification that was not in the search parameters. Other ways to identify a spectrum Database searching is not the only strategy. The main alternatives differ in what prior knowledge they need, and therefore in how they interact with a customised database, as Table 1 summarises. Table 1. Principal strategies for identifying peptides from tandem mass spectra. Strategy Prior knowledge required Strength Main limitation for proteogenomics Sequence database search Protein sequence database Sensitive, well-understood statistics Cannot find sequences absent from the database Spectral library search Library of previously identified spectra Fast, highly specific Library rarely contains novel peptides Open or mass-tolerant search Database plus wide precursor tolerance Reveals unexpected modifications Larger search space, harder error control De novo sequencing None beyond fragmentation rules Can propose sequences absent from any database Ambiguities and incomplete spectra limit accuracy Peptide-centric search A specific candidate peptide Direct test of one hypothesis against all alternatives Answers only the question asked Spectral library searching compares an observed spectrum with a library of previously identified experimental spectra and is fast and specific, but by construction a library contains only what has already been seen. Predicted spectral libraries partly remove that constraint. Prosit, described by Gessulat and colleagues in 2019, is a deep learning model trained on very large sets of synthetic peptide spectra that predicts fragment intensities and chromatographic retention times for arbitrary sequences, and similar models have followed. With such predictions, a novel peptide from a proteogenomic database can be given an expected spectrum and an expected elution time, both of which become powerful validation tools. De novo sequencing reads the sequence directly from the mass ladder. It needs no database and so is in principle ideal for finding the unexpected, and deep learning approaches have raised its accuracy considerably. In practice, full-length correct de novo sequences are obtained for a minority of spectra, and the method is most useful in proteogenomics as a complement: for proposing candidates in samples without matched sequencing data, or for checking whether a spectrum matched to a novel database entry independently supports that sequence. Peptide-centric searching inverts the usual question. Instead of asking which peptide best explains each spectrum, it asks, for one candidate peptide, whether any spectrum supports it better than every alternative explanation. Chapter 9 discusses PepQuery, the best-known implementation, as a validation step. The target-decoy approach A search engine will always report a best match, whether or not the true peptide is in the database. The central problem of peptide identification is therefore distinguishing correct from incorrect PSMs. A score threshold alone does not solve it, because the distribution of scores for incorrect matches depends on the data, the instrument, the database and the parameters. The solution that became standard is the target-decoy strategy, formalised by Elias and Gygi in 2007. Alongside the real protein sequences, the targets, the search includes an equal number of decoy sequences that cannot be correct: typically the target sequences reversed, or shuffled in a way that preserves amino-acid composition and, often, the enzymatic cleavage sites. Every spectrum is searched against the combined database. Because a spectrum that does not truly belong to any target peptide is, by assumption, equally likely to be matched by chance to a target or a decoy, the number of decoy matches above any score threshold estimates the number of incorrect target matches above that threshold. The false discovery rate is then estimated as the ratio of decoys to targets passing the threshold, and the threshold is chosen so that this ratio is at most, say, 1 percent. The method is simple, general and empirical, which is why it prevailed. But it rests on assumptions that proteogenomic databases can quietly violate. Decoys must be generated from the same database as the targets, with the same size and composition, so that random matches fall on them at the same rate. Every spectrum must be given the same chance to match a decoy as a target. And the false discovery rate it estimates is a property of the whole set of matches passing the threshold, not of any particular match. A 1 percent rate across a list of fifty thousand PSMs says that roughly five hundred are wrong; it says nothing about which ones. If the wrong ones are concentrated in a small subgroup, such as the novel peptides that are the reason for the experiment, the error rate within that subgroup can be very much higher than 1 percent. Chapter 4 is devoted to that problem. Rescoring and machine learning A single search-engine score uses only part of the information in a match. Other features also distinguish correct from incorrect PSMs: the precursor mass error, the number of matched fragment ions, the fraction of total intensity explained, the charge state, the number of missed cleavages, and the difference in score between the best and second-best candidate. Percolator, introduced by Käll and colleagues in 2007, uses a semi-supervised support vector machine that learns, iteratively from the target and decoy matches of each dataset, how to combine such features into a better discriminant score. It typically increases the number of identifications at a fixed false discovery rate. More recent rescoring adds features derived from predicted spectra and retention times: how well the observed fragment intensities agree with those predicted by a model such as Prosit for the candidate sequence, and how close the observed elution time lies to the predicted one. These features are especially valuable for peptides with unusual sequences, including non-tryptic peptides eluted from human leukocyte antigen molecules, where classic scores perform poorly. They carry a caution, however. Rescoring learns from the data what a correct match looks like, and if the training set is dominated by canonical peptides, novel peptides that behave differently may be treated unfairly in either direction. The safest practice is to check that rescoring behaves sensibly on the novel subset specifically, not only on the whole. From peptides to proteins The final step, protein inference, assembles peptide identifications into protein identifications. It is harder than it looks, because many peptides are shared between proteins: between isoforms of one gene, between members of a gene family, and, in proteogenomic databases, between a reference protein and its variant or alternative version. A peptide that is shared cannot by itself indicate which of its possible parent proteins was present. Protein inference algorithms apply parsimony, reporting the minimal set of proteins that explains the observed peptides, or probabilistic models that assign each protein a probability given the peptides. For proteogenomics, the lesson is that evidence for a novel product nearly always rests on peptides unique to that product: the variant-containing peptide, the novel junction peptide, the peptide from the unannotated reading frame. Shared peptides say that something from the locus was present, not that the novel form was. Because unique peptides are few, many novel claims rest on one or two PSMs, which is exactly why they require the extra scrutiny the rest of this book describes. What the database controls It is now possible to state precisely what the database does. It determines which sequences can be identified at all. It determines how many candidates each spectrum competes against, and so the distribution of chance scores and the threshold needed to control errors. It determines the decoys, and so the error estimate. It determines which peptides are unique and which are shared, and so what protein-level conclusions can be drawn. A proteogenomic analysis is, in effect, an experiment designed by building a database, and the rest of the pipeline reports the outcome. The next chapter examines how such databases are made. Chapter 3: Building the Customised Database A customised protein sequence database is assembled from pieces, each derived from a different kind of evidence and each translating nucleotide information into amino-acid sequences in its own way. The choices made here settle most of what the downstream search can and cannot find. This chapter goes through the main sources in rough order of how directly they reflect the sample: from generic translations of the genome, through population and sample-specific variants, to transcripts assembled from the sample's RNA and reading frames observed on its ribosomes. The spectrum of databases Proteogenomic databases differ in two properties that pull against each other: breadth, meaning how much of the possible proteome they include, and specificity, meaning how much of their content is actually likely to be present in the sample. As Table 2 sets out, the generic approaches maximise breadth at great cost in size, while sample-specific approaches keep the database small but can include only what the sequencing data detected. Table 2. Main sources of sequences for proteogenomic databases. Source Input data What it adds Relative size Main risk Six-frame genome translation Genome sequence Every possible reading frame Very large Severe loss of sensitivity, many false matches Three-frame transcript translation Annotated or assembled transcripts Alternative frames and untranslated regions Large Most frames are not translated Variant database DNA or RNA variant calls, or public catalogues Single amino-acid variants, small indels Small to moderate Mis-called variants, variants in unexpressed alleles Splice-junction database RNA-seq junction reads Novel exon-exon junction peptides Small to moderate Alignment artefacts, junctions from immature transcripts Transcript assembly RNA-seq reads Novel isoforms, fusions, unannotated genes Moderate Assembly errors, arbitrary reading-frame choice Ribo-seq ORFs Ribosome profiling Translated ORFs including short and upstream ORFs Small to moderate Ribosome occupancy without stable product Genome translation The oldest strategy is to translate the entire genome in all six reading frames, three on each strand, splitting at stop codons, and to search against every resulting stretch of amino acids above some minimum length. It requires no annotation and makes no assumptions, which is why it served well for early annotation of small genomes. For a human-sized genome it is, however, a poor default. The six-frame translation of the human genome is vastly larger than the reference proteome, and nearly all of it is never translated. Every spectrum faces orders of magnitude more candidates, the score needed to reach a given error rate rises sharply, and the identification rate falls. Worse, the novel matches that do pass are drawn from a region of search space where the prior probability of a real peptide is tiny, so a larger fraction of them are wrong than the global threshold implies. Six-frame searches of human data can still be useful for specific questions, and in a restricted form, for instance translating only regions with evidence of transcription, but they should not be the starting point. Three-frame translation of transcripts is a middle ground. Here each annotated or assembled transcript is translated in its three forward frames, so that untranslated regions, retained introns and alternative frames are included, but only for sequences known to be transcribed. This yields a database much smaller than a six-frame genome translation but still much larger than the proteome, and like the six-frame approach it is dominated by sequences that are not translated. Variant databases Most proteogenomic studies of human samples add sequence variants. The inputs are variant calls in a standard format, derived from whole-genome, whole-exome or RNA sequencing of the sample, or from public catalogues of population variation such as dbSNP and gnomAD, or of cancer mutations such as COSMIC. A tool then annotates each variant against the reference gene models, determines its effect on the coding sequence and writes out the altered protein sequences. The details matter more than they might seem. A missense variant changes one residue, and the affected tryptic peptide is the unit that matters, so many tools write out just that peptide, with flanking context, rather than duplicating the whole protein. A variant that creates or destroys a lysine or arginine changes where trypsin cuts, altering not one peptide but the boundaries of two. Insertions, deletions and frameshifts must be handled by translating the altered coding sequence through to the next stop codon. Variants that lie close together on the same haplotype should in principle be applied together, since a peptide carrying both exists only if they are on the same chromosome. And a variant falling in multiple transcripts of the same gene may produce different peptides in each. Several well-established tools perform these steps. customProDB, an R package described by Wang and Zhang in 2013, generates customised databases from RNA-seq data, including variants, novel junctions and a filter that keeps only transcripts expressed above a threshold in the sample. PGA, a Bioconductor toolkit from the same group, extends this to include novel transcripts and to handle the downstream search and error control. Other pipelines exist in Galaxy-based platforms, in Python packages that operate on variant call files, and within larger frameworks for neoantigen discovery. A basic question is which variants to include. Adding all known population variants, many millions of them, inflates the database with sequences mostly absent from any one sample. Sample-specific variants are far fewer, but their inclusion depends on having good sequencing data and on the variant caller's error rate. A reasonable general rule is to include the variants called in the sample itself, with quality filters, and to treat population catalogue variants as a separate, lower-prior category if they are used at all. Splice junctions and assembled transcripts RNA sequencing reads that span exon-exon junctions provide direct evidence of which splice events occurred. After aligning reads with a splice-aware aligner such as STAR or HISAT2, junctions supported by a minimum number of reads, and ideally not previously annotated, can be translated into short peptide sequences covering the junction. Sheynkman and colleagues showed in 2013 that searching such junction peptides allows novel splice events to be confirmed at the protein level, though with modest yield. Because each junction contributes only one or a few peptides, the database grows little, and the added sequences are specific to what was observed in the sample. The reading frame is the subtle part. A novel junction between two annotated coding exons can be translated in the frame of the upstream exon, and the result is unambiguous if the downstream frame is preserved. If the junction shifts the frame, or joins exons not normally connected, the translation beyond the junction may be a novel sequence. Tools typically generate all frames consistent with an upstream open reading frame, or, more conservatively, only the frame of the annotated upstream coding sequence. Full transcript assembly, with tools such as StringTie or Trinity, goes further. It reconstructs entire transcripts from the reads, including isoforms that combine known exons in new ways, transcripts from unannotated loci and fusion transcripts. The assembled transcripts must then be translated, which requires choosing reading frames. Programs such as TransDecoder predict the most likely coding region of each transcript, usually based on length and sequence composition. That choice is a hypothesis in its own right: a transcript may contain several plausible reading frames, or none that is actually translated. Assembly also produces errors, particularly for lowly expressed transcripts and in repetitive regions, and chimeric assemblies can generate peptide sequences that exist nowhere in the cell. Assembled databases are therefore powerful for discovering isoforms and fusions, and they require the same scepticism as any other computational prediction. Fusion genes deserve special mention. Dedicated fusion callers, such as STAR-Fusion or Arriba, identify chimeric transcripts from RNA-seq with higher specificity than general assembly. The peptide spanning the fusion junction, where it maintains or shifts the reading frame of the downstream partner, is a sequence unique to the tumour, and in some cancers, such as those driven by recurrent fusions, a highly attractive target. Ribosome profiling Ribosome profiling, introduced by Ingolia and colleagues in 2009, sequences the fragments of messenger RNA protected by translating ribosomes from nuclease digestion. These footprints, around 28 to 30 nucleotides long, map to the positions ribosomes occupied at the moment of harvest. Because ribosomes move three nucleotides at a time, the footprints show a characteristic three-nucleotide periodicity in actively translated regions, and this periodicity can be used to infer which reading frame is being translated. Treatment with drugs that stall ribosomes at initiation sites, such as harringtonine or lactimidomycin, further allows start codons to be mapped. Algorithms such as RiboTaper, ORF-RATER, RibORF and PRICE use these signals to call translated open reading frames across the transcriptome. Their output is exactly what a proteogenomic database wants: a list of reading frames with direct evidence of translation, including many that annotation ignores, such as upstream open reading frames, short reading frames in long non-coding RNAs and reading frames overlapping known genes in alternative frames. A Ribo-seq-derived database is much smaller and much more specific than a three-frame translation, because it contains only frames that ribosomes were seen to engage. Ribosome profiling is not a direct measurement of protein, however. Ribosome occupancy shows that translation began, not that a stable product accumulated. Many of the products of short reading frames are expected to be rapidly degraded, and some may never be detectable by standard shotgun proteomics even when translation is real. Matched Ribo-seq and mass spectrometry data from the same samples is valuable for this reason: it allows a translated frame to be recorded as translated, and its product, if found, to be recorded as detected, without conflating the two. Assembly and housekeeping The final database is assembled from the reference proteome plus the chosen additions, and then cleaned. Several housekeeping steps matter a great deal. Redundancy should be removed. Many novel sequences share most of their residues with reference proteins; a variant protein differs by one residue, a junction isoform by a few. Including full redundant sequences inflates the database and complicates protein inference without adding identifiable peptides. Many pipelines therefore write out only the novel peptides, or short sequences around the novel region, and remove any novel peptide whose sequence already exists elsewhere in the reference, including after treating leucine and isoleucine as equivalent, since mass spectrometry cannot distinguish them. Every entry should carry provenance. The sequence header should record which source generated it, which gene and transcript it derives from, which variant or junction it covers, and what supporting evidence existed, for instance the number of RNA reads or the Ribo-seq score. Without provenance, the novel identifications at the end cannot be traced back or stratified, and the separate error control recommended in the next chapter becomes impossible. Common contaminants should be included: keratins, trypsin, serum albumin and other proteins routinely introduced during sample handling. A contaminant spectrum with no correct target may otherwise be matched to a novel sequence. Decoys should be generated from the final combined database, not from the reference alone, and with the same method for all parts. If decoys are made only from the reference, the novel sequences have no decoy counterparts and their false matches go uncounted. A practical design sequence In practice, a careful database design for a human proteogenomic study might proceed as follows. Start from a well-curated reference proteome, such as the GENCODE basic set or UniProt reviewed human proteins with isoforms, plus contaminants. Add sample-specific variants called from matched DNA or RNA sequencing data, filtered for quality and, for RNA-based calls, for adequate read support. Add novel junction peptides supported by several RNA reads. If fusions matter, add junction peptides from a dedicated fusion caller. If non-canonical translation is the question and Ribo-seq data exist, add Ribo-seq-supported reading frames; if not, add a restricted three-frame translation of expressed transcripts, recognising the cost. Remove redundant sequences, record provenance, generate decoys from the whole, and keep the categories labelled so that they can be evaluated separately after the search. The principle behind this sequence is to include sequences in proportion to the evidence that they might be present. Each addition should be justified by data from the sample or by a specific question. Everything included is a candidate for identification, but also a candidate for a false match, and the next chapter explains why that second fact becomes more pressing as the database grows. Hashtags: #Proteogenomics #Genomics #Transcriptomics #MassSpectrometry #ShotgunProteomics #CustomProteinDatabases #SampleSpecificDatabases #VariantPeptides #SomaticMutations #AlternativeSplicing #SpliceJunctionPeptides #NonCanonicalTranslation #RibosomeProfiling #NeoantigenDiscovery #Immunopeptidomics #PeptideSpectrumMatching #TargetDecoyStrategy #FalseDiscoveryRate #ProteinInference #SpectralValidation #SyntheticPeptideValidation #RetentionTimePrediction #CPTAC #PersonalizedProteomics #FutureOfProteogenomics

Latest Book Releases:

WELCOME TO THE INTERNATIONAL STUDENTS LIBRARY

bottom of page