Welcome to the VBNN Digital Library
Unlock a Vast Knowledge Ecosystem
Featuring over 30,000 books, academic papers, illustrations, and expert insights—continuously updated to support your research and professional growth.
Welcome to our library!
Here, you will find an exclusive collection created 100% by our own faculty, meaning you will not find these resources anywhere else. Over the last 20 years, our team has written much more than what is currently online, and we are actively working to upload our complete back catalog. We update our platform regularly, so be sure to check back from time to time. If you ever need help finding a specific resource, you can always contact us!
Maximize Your Access
Log in to instantly view and download tailored resources directly aligned with your specific program and curriculum.
Ready to begin? Sign in above to explore your personalized dashboard.
Please note: Login is only possible using your institutional email address; otherwise, the system will not recognize your account.
VBNN Library AI
Introducing our fully integrated Library AI. Designed to support your research, you may submit inquiries in any language and receive precise, evidence-based responses drawn exclusively from our published scholarly articles and textbooks.
Search...
Latest Publications:
Search this site
Results found for empty search
- The Business of Space Tourism (Liability, Market Demand, and Infrastructure)
Download the Book (PDF): Introduction On 22 January 2026 six paying passengers climbed into a Blue Origin capsule in West Texas, rode a New Shepard booster above the Kármán line, floated for a few minutes, and parachuted back into the desert. It was the thirty-eighth flight of the system and, as it turned out, the last for some time. Eight days later Blue Origin announced that New Shepard would stop flying for "no less than two years" so that the company could move people and money onto its lunar lander. Virgin Galactic, the other American seller of suborbital seats, had not flown a customer since June 2024; its new Delta-class spaceships, first promised for 2026, were by August 2026 scheduled to debut in February 2027. For most of 2026, then, the world's two established suborbital tourism businesses were simultaneously out of service. Anyone who wanted to buy a ride to space in the autumn of 2026 could put down a deposit, but could not fly. This was not a story of collapse. Blue Origin had carried by its own count 98 people above the Kármán line, and it stopped not because customers vanished but because its owner found a better use for its engineers. Virgin Galactic had sold its latest tranche of seats at 750,000 dollars each, the highest price it had ever charged, and reported that the tranche sold out ahead of schedule. Orbital demand was, if anything, stronger: private astronauts flew to the International Space Station on SpaceX Dragon capsules, a cryptocurrency entrepreneur chartered a Dragon to circle the poles in 2025, and NASA handed out new private-mission awards in early 2026 to both Axiom Space and Vast. The people were there. The money was there. What was missing was something steadier: a reason to believe that a passenger spaceflight business could run the way a business runs, day after day, rather than as a sequence of spectacular one-off events. That gap between demand and operation is the subject of this booklet. Passenger spaceflight has spent a quarter-century being discussed as if its central question were whether rich people would pay to go to space. That question was answered long ago, and the answer is yes. The harder questions are about what sort of enterprise can turn that willingness into a durable stream of profitable flights, and what the law, the insurance market and the public purse have to look like for that enterprise to exist. Those questions sit at the intersection of three subjects that are usually treated separately: the economics of demand and pricing, the law of liability and safety regulation, and the physical and financial infrastructure of spaceports. This booklet treats them together because they cannot be understood apart. The argument The argument can be put in a sentence. Passenger spaceflight becomes a viable business only when it reaches a high, reliable flight rate at low marginal cost, and nearly every arrangement that currently surrounds it, from ticket pricing and customer waitlists to the federal "learning period" on safety rules and the financing of public spaceports, was designed for a low-cadence luxury experiment rather than for that high-cadence operation. Each arrangement made sense when it was adopted. Taken together they now form a kind of equilibrium in which operators can raise prices, delay service and pause flights without much penalty, while the risks of those choices fall on deposit holders, on local taxpayers who built runways, and on a regulator that has deliberately limited its own authority. Flight rate is the hinge because every other variable in the business moves with it. A suborbital spaceship that flies once a month must recover its enormous fixed costs from twelve flights a year; the same ship flying twice a week spreads them over a hundred. Ticket prices that look extravagant at low cadence become ordinary when the fixed costs are amortised more widely, and only then does the deeper, more price-sensitive layer of demand that surveys have detected for two decades come into reach. The safety case changes with flight rate too: a vehicle with ten flights has almost no statistical record, while a vehicle with five hundred begins to produce the kind of data from which sensible rules can be written. Spaceports, finally, earn their keep through launch fees and associated activity, so a spaceport whose anchor tenant flies rarely is a public cost centre, while one whose tenant flies often can become a genuine piece of transport infrastructure. What this booklet covers The chapters move from the product outward. Chapter 1 defines what exactly is being sold, because "space tourism" covers experiences that differ by three orders of magnitude in price, from a proposed stratospheric balloon ride to a multi-week stay on an orbiting station. Chapter 2 tells the history of the suborbital race as a history of development risk: why companies founded around 2000 took two decades to carry their first paying customers, and what accidents, bankruptcies and schedule slips reveal about the true cost of building a passenger spacecraft. Chapter 3 turns to demand and pricing, examining the market studies that have shaped expectations since 2002 and asking what can honestly be said about the price elasticity of a product sold to a few hundred people a year. Chapter 4 builds the unit economics of a suborbital operation and shows why flight rate dominates every other variable, using the figures Virgin Galactic has itself disclosed for its Delta fleet. The middle of the booklet turns to law. Chapter 5 examines the liability architecture of American commercial human spaceflight: informed consent, reciprocal waivers of claims, state immunity statutes, and the federal backstop for damage to third parties on the ground. Chapter 6 examines regulation proper, especially the congressional moratorium on occupant-safety rules known as the learning period, the 2025 recommendations of the FAA's rulemaking committee, the August 2025 executive order on commercial space, and the different approaches taken in the United Kingdom and Europe. Chapter 7 turns to infrastructure and public finance, using the experience of Spaceport America, Spaceport Camden, Spaceport Cornwall and others to ask when public investment in spaceports pays off and when it becomes a stranded asset. Chapter 8 considers orbital tourism, the segment where prices run to tens of millions of dollars per seat and where the retirement of the International Space Station, planned for 2030, is forcing NASA to decide whether private visitors will anchor the next generation of stations. The conclusion draws these threads together into an argument about what a mature passenger spaceflight industry would require, and who should bear the cost of getting there. A note on the moment This field moves quickly, and any account of it risks being overtaken. The booklet is written from the position of late 2026: Virgin Galactic preparing its Delta ships for flight test, New Shepard grounded by choice, the FAA's learning period scheduled to expire on 1 January 2028 unless Congress extends it again, the International Space Station with four years of planned life remaining, and commercial space stations still on the drawing board or in the cleanroom. Some specific figures quoted here will change by the time they are read. The structural argument should not. Whatever vehicle flies next, and whichever company flies it, the relationships between cadence, price, liability and infrastructure described here will decide whether passenger spaceflight becomes an industry or remains an occasional spectacle. The approach throughout is that of an analyst rather than an enthusiast or a sceptic. Space tourism attracts both in abundance. Enthusiasts tend to treat each new vehicle announcement as proof that the market is about to open, and each price increase as evidence of irresistible demand. Sceptics tend to treat the whole enterprise as a vanity project for billionaires, to be judged mainly by its carbon footprint or its social optics. Neither view helps anyone decide whether to invest in a spaceship company, approve a spaceport bond, draft a liability statute or sign a passenger waiver. Those decisions require an understanding of how the business actually works: where the money comes from, where the risks sit, and which constraints are physical and which merely institutional. That is what this booklet sets out to provide. Chapter 1: What Is Being Sold "Space tourism" is a phrase that hides more than it reveals. It is applied with equal confidence to a proposed balloon ride to the stratosphere, a ten-minute hop in a capsule over West Texas, a three-day orbital cruise in a SpaceX Dragon, and a two-week visit to the International Space Station. These experiences differ in price by roughly two orders of magnitude, in physical risk by a similar margin, in regulatory treatment, in the kind of customer who buys them, and in the infrastructure they need. A business analysis that lumps them together will reach conclusions that are true of none of them. The first task, then, is to take the phrase apart. The most useful way to do so is to ask what the customer actually receives. Every product in this market sells some combination of four things: altitude, which determines whether the passenger sees the black sky and the curvature of the Earth; weightlessness, measured in minutes for suborbital flights and in days for orbital ones; the status of having "been to space", which depends heavily on where that boundary is drawn; and a surrounding experience of training, ceremony and community that is often worth as much to the buyer as the flight itself. Different vehicles bundle these differently, and their prices reflect those bundles more than their engineering costs. A ladder of products At the bottom of the ladder, in price if not in ambition, are stratospheric balloon flights. The concept is simple: a pressurised cabin carried by a very large balloon to around 30 kilometres, high enough to see a dark sky and the curve of the horizon, followed by a slow descent. There is no weightlessness and no plausible claim to have reached space. The most prominent American venture, Space Perspective, sold seats at 125,000 dollars and flew an uncrewed test of its capsule in September 2024. By early 2025 it had run out of money, furloughed its staff and effectively ceased operations, having never carried a paying passenger. Its failure is instructive precisely because its technology was the least demanding in the field. Even a product that avoids rockets entirely needs years of pre-revenue development, and investors' patience for that runway is finite. The next rung is suborbital rocket flight, which is what most people mean by space tourism. Here the passenger rides a rocket-powered vehicle on a steep arc above an altitude that some authority recognises as the edge of space, experiences three or four minutes of weightlessness at the top, and returns to the ground the same day. Two American designs have carried paying customers. Virgin Galactic's SpaceShipTwo is a winged rocket plane released from a carrier aircraft at about 15 kilometres altitude; it fires its hybrid rocket motor, climbs to between 80 and 90 kilometres, and glides back to a runway. Blue Origin's New Shepard is a vertical system: a reusable booster lifts a six-seat capsule above 100 kilometres, the capsule separates, and after the coast it descends under parachutes while the booster lands itself on a pad nearby. A Chinese company, Deep Blue Aerospace, began selling seats in late 2024 for a suborbital service it hopes to start in 2027, pricing them at 1.5 million yuan, about 210,000 dollars. Above suborbital flight lies orbit, and here the price leaps by roughly a factor of fifty. To stay in orbit a spacecraft must reach about 7.8 kilometres per second, which requires on the order of thirty times the energy per kilogram needed simply to coast up to 100 kilometres and fall back, a large orbital rocket, and a spacecraft able to survive a high-energy re-entry. Orbital tourism has itself split into two products. One is the free-flying charter, in which a private customer buys the whole of a SpaceX Crew Dragon mission and flies it without visiting any station: Inspiration4 in 2021, Polaris Dawn in 2024, which included the first spacewalk by private astronauts, and Fram2 in 2025, which carried four first-time fliers over both poles. The other is the station visit, in which a customer buys a seat on a mission to the International Space Station. Between 2001 and 2009 seven private individuals flew to the station on Russian Soyuz spacecraft through the broker Space Adventures, beginning with Dennis Tito, whose ticket was widely reported at about 20 million dollars. Since 2022 Axiom Space has organised such visits using Dragon, with seats reported at between 50 and 60 million dollars. At the top of the ladder, still hypothetical, is travel beyond Earth orbit. The Japanese businessman Yusaku Maezawa bought a circumlunar flight on SpaceX's Starship in 2018, branded it dearMoon and recruited a crew of artists; in June 2024 he cancelled the project, citing the uncertain schedule of the vehicle. No other lunar tourism mission is under contract as of 2026. One further product is often mentioned alongside tourism though it is really a different business: point-to-point suborbital transport, in which a rocket would carry passengers between distant cities in under an hour. SpaceX has promoted the idea for its Starship vehicle, and in 2019 UBS estimated that such travel could eventually become a market of about 20 billion dollars a year. No vehicle is being built or certified for that purpose, and the obstacles, including noise, launch hazard areas near cities, passenger acceleration loads and the need for aviation-grade reliability, are formidable. It belongs in any complete picture of the field as a long-term possibility that shapes investor expectations. It is not a product anyone can buy, and it plays no part in the business analysis that follows, except as a reminder that the ultimate prize in this industry has always been imagined as transport rather than tourism. Table 1 sets out the main products as they stood in September 2026. The price column deserves caution. Only Virgin Galactic and Deep Blue publish per-seat prices; the others are negotiated privately and reported by journalists, or not at all. Table 1. The passenger spaceflight product ladder, September 2026. Product Operator Altitude Zero-g Seat price Status Stratospheric balloon Space Perspective About 30 km None 125,000 dollars (advertised) Company ceased operations in 2025 without flying passengers Suborbital spaceplane Virgin Galactic Delta class About 80 to 90 km A few minutes 750,000 dollars (latest tranche) First Delta flight scheduled for February 2027 Suborbital capsule Blue Origin New Shepard Above 100 km A few minutes Not published Paused from January 2026 for at least two years Suborbital capsule (China) Deep Blue Aerospace About 100 km (planned) A few minutes (planned) 1.5 million yuan (about 210,000 dollars) Tickets sold; service targeted for 2027 Orbital free-flyer SpaceX Crew Dragon charter Low Earth orbit, several hundred km and above Several days Whole-vehicle charter, price not published Flown in 2021, 2024 and 2025 Station visit Axiom Space via Crew Dragon ISS, about 400 km One to three weeks Reported 50 to 60 million dollars Fifth mission awarded for early 2027 Sources: company announcements; Virgin Galactic shareholder communications (2026); Blue Origin (January 2026); SpaceNews; South China Morning Post (2024); New Space Economy market analysis (March 2026). Where space begins, and why it matters commercially The boundary of space is not a physical surface but a convention, and the choice of convention has commercial consequences. The Fédération Aéronautique Internationale, which keeps aviation and astronautics records, uses 100 kilometres, the so-called Kármán line. The United States Air Force and, for its now-retired astronaut wings programme, the Federal Aviation Administration used 50 miles, about 80 kilometres. Virgin Galactic's spaceplane flies to between 80 and 90 kilometres, above the American line and below the international one. New Shepard flies above 100 kilometres and Blue Origin has made a point of saying so. For the buyer, the difference between 86 and 106 kilometres is almost invisible. The sky is black at both altitudes, the curvature of the Earth is obvious, and weightlessness lasts a similar few minutes. But the status good being sold, the ability to say one has been to space, depends on the definition, and marketing departments know it. The FAA ended its Commercial Space Astronaut Wings programme at the end of 2021, precisely because the number of people reaching space on commercial vehicles had grown too large for an individual award to mean much. It now simply lists on its website the people who have flown above 50 miles on FAA-licensed vehicles. That decision quietly acknowledged something important: once flights become routine, the badge loses scarcity value, and the business must stand on the experience itself. Who buys, and why it shapes the product The customer base for these products is small and unusual. Virgin Galactic reported in 2026 that about 675 people held reservations, drawn from dozens of countries, many of whom had signed up years earlier at lower prices. The most detailed early market research, a 2002 survey conducted by the consulting firm Futron with the pollster Zogby, interviewed 450 people with annual incomes of at least 250,000 dollars or net worth of at least one million dollars. The customers who have actually flown are more concentrated still: technology and finance founders, celebrities, a handful of scientists and educators flown on sponsored seats, and a surprising number of people who won or were given their seats by someone else. Three features of this customer base shape the product. First, the buyers are extraordinarily wealthy, which means that price is rarely the binding constraint for the first few hundred sales but becomes decisive very quickly after that; this is the subject of Chapter 3. Second, many buyers are purchasing a story as much as a ride, and the story has a social dimension. Flights are sold as small cohorts with shared training, and operators invest heavily in the ceremony around launch: the pre-flight dinners, the family viewing areas, the post-flight wings. Third, a meaningful share of seats is bought not by individuals but by institutions: governments seeking a national first, universities and research agencies flying experiments with a human operator aboard, and brands buying publicity. Axiom's 2025 Ax-4 mission, which launched on 25 June 2025, is a clear example. Three of its four seats carried government-sponsored astronauts from India, Poland and Hungary, the first from those countries to fly in decades or ever. That is a tourism flight in its mechanics and a national space programme in its purpose. The institutional buyer matters more than its numbers suggest, because it is less price-sensitive and more repeatable than the thrill-seeking individual. A billionaire flies once; a national space agency that has flown one astronaut may well want to fly another. Virgin Galactic has reported growing interest from research institutions in multi-seat bookings for its Delta ships, and the company has long flown experiments for research customers. Chapter 8 will show that in orbit, government-sponsored visitors have become the backbone of demand. The experience and the envelope Finally, it is worth dwelling on the fact that the flight is only part of what is sold. A Virgin Galactic customer has historically received several days of preparation at Spaceport America in New Mexico, including training, medical checks and social events, as well as membership of a community of "future astronauts" that the company cultivates for years before the flight. Blue Origin's customers train for about two days in Texas. Orbital customers train for months, both because the flight is longer and more hazardous and because they must learn to live and work in a spacecraft. For an orbital visitor the training is itself a substantial part of the experience, and for an agency-sponsored astronaut it is a part of the national investment. This envelope has two business implications. It adds to cost, because training facilities, staff and hospitality are not trivial, and it adds to value, because it extends the customer's engagement and gives the operator room to differentiate. But it also shapes the liability position discussed in Chapter 5. The more an operator trains, briefs and prepares its passengers, the stronger its argument that they gave genuinely informed consent to the risks they ran. The ceremony is also a legal instrument. Who can physically go A product's market is bounded not only by who can afford it but by who can physically consume it, and here too the products on the ladder differ sharply. A suborbital flight subjects its passengers to several times the force of gravity during the climb and again during re-entry, for periods measured in seconds to a minute or two, followed by a few minutes of weightlessness. Healthy people tolerate these loads well, and the experience of the last five years has shown that age alone is not disqualifying. Wally Funk flew on New Shepard at 82, William Shatner at 90, and in 2024 Ed Dwight, who in the early 1960s had been the first Black trainee in the Air Force test pilot programme from which astronauts were then drawn but was never selected to fly, went to space on New Shepard at 90. The operators screen their customers medically, but they do so under their own policies: American law, as Chapter 6 explains, imposes no medical standard on passengers. An orbital flight is a different physical proposition. Days or weeks of weightlessness bring space motion sickness, fluid shifts towards the head, disturbed sleep and the practical challenges of living in a small cabin, as well as the loads of launch and re-entry and a longer exposure to radiation. Orbital customers undergo longer medical assessment and months of training, and those visiting the International Space Station must meet medical standards set by NASA and its partners. The population able and willing to go through that process is much smaller than the population able to take a ten-minute ride. This physical envelope has a direct commercial consequence. The Futron study discussed in Chapter 3 found that relaxing its assumed fitness requirements, to include people of average fitness under sixty-five, roughly doubled its forecast of suborbital demand. The operators' growing record of carrying older and less athletic passengers without incident is therefore not merely a human-interest story. It is evidence that the addressable market is larger than the most cautious early studies assumed, and every flight that carries an older passenger safely expands it a little further. The same record, of course, would be scrutinised closely after any accident in which a passenger's health played a part. What the ladder tells us Laid out this way, the product ladder shows why generalisations about "the space tourism market" are so often wrong. The suborbital market is a market in which the product is brief, the price is in the hundreds of thousands of dollars, the buyer is an individual or a research group, and viability depends on flying often from dedicated ground infrastructure. The orbital market is a market in which the product lasts days or weeks, the price is in the tens of millions, the buyer is increasingly a government, and viability depends on access to a very small number of rockets, capsules and stations controlled by a handful of firms and by NASA. The stratospheric balloon market, which should have been the easiest to open, has so far been the first to fail. What unites these products is not price, customer or technology. It is that each of them has so far been produced at very low volume, and each business model assumes that volume will rise. The chapters that follow examine what stands between the flights that have happened and the cadence those models require. Chapter 2: The Long Road to a Ticket On 4 October 2004 a small winged rocket called SpaceShipOne climbed above 100 kilometres for the second time in five days and won the Ansari X Prize, a ten-million-dollar award for the first privately built vehicle to carry a person to space twice within two weeks. The vehicle had been designed by Burt Rutan's company Scaled Composites and financed by the Microsoft co-founder Paul Allen. Within days Richard Branson announced that his new company, Virgin Galactic, would license the technology and begin carrying tourists on a larger successor. The early promises suggested commercial flights within a few years. The first Virgin Galactic flight with private paying customers took place in August 2023, nearly nineteen years later. That interval is the central fact of the industry's history, and it is usually told as a story of hubris or bad luck. It is better understood as a lesson in the economics of development risk. Building a vehicle that carries untrained members of the public on a rocket, and then flying it repeatedly, turned out to be a much larger engineering and organisational task than anyone priced in. The delays were not an accident of one company's management. They were the market's discovery of the true cost of the product. This chapter traces that discovery and draws out what it implies for anyone valuing a passenger spaceflight business today. The Virgin Galactic path The history of SpaceShipTwo is punctuated by the kinds of events that define development risk: accidents, redesigns and long groundings. In July 2007, during a ground test at the Mojave Air and Space Port in California, a tank of nitrous oxide exploded and killed three Scaled Composites employees. The oxidiser was part of the hybrid propulsion system inherited from SpaceShipOne, and the accident forced a reassessment of how that system was handled. Powered flight testing of the first SpaceShipTwo, VSS Enterprise, did not begin until 2013. On 31 October 2014, during a powered test flight over the Mojave Desert, VSS Enterprise broke apart. The co-pilot, Michael Alsbury, was killed; the pilot, Peter Siebold, survived after parachuting from the disintegrating vehicle. The National Transportation Safety Board, which led the investigation, found that the co-pilot had unlocked the vehicle's "feather" re-entry system early, during the transonic phase of the climb, and that aerodynamic loads then deployed it without any further command. But the Board did not stop at pilot error. Its probable cause statement placed the weight on Scaled Composites' failure to consider and protect against the possibility that a single human error could cause a catastrophic hazard, and it criticised the FAA for approving waivers of its own hazard-analysis requirements without adequately examining the design. The accident became, and remains, the defining safety event in American commercial human spaceflight. Chapter 6 returns to what the regulator learned from it. Virgin Galactic brought a second vehicle, VSS Unity, into flight test in 2016 with a mechanical inhibit added to the feather system. Unity first reached the American 50-mile boundary of space in December 2018 and first carried a passenger other than its pilots, the company's chief astronaut instructor, in February 2019. In the same year the company became the first human spaceflight firm to list on a public stock exchange, through a merger with a special-purpose acquisition company. The listing gave it capital and visibility but also subjected every schedule slip to quarterly scrutiny. On 11 July 2021 Unity carried Richard Branson and three company employees to space, nine days before Jeff Bezos flew on New Shepard. The flight was a public-relations triumph, but it produced an unexpected coda. Weeks later it emerged that the vehicle had strayed outside its assigned airspace during descent, and the FAA prohibited further Virgin Galactic flights until the company addressed the deviation and its reporting of it. The episode was resolved within weeks, but it showed how a vehicle flying to space from inland airspace is embedded in the ordinary machinery of air traffic control, and how regulatory authority over the public's safety persists even where authority over passengers' safety is limited. The company then took Unity and its carrier aircraft out of service for nearly two years of refurbishment and upgrade. Commercial service finally began on 29 June 2023 with a research flight for the Italian Air Force, followed in August 2023 by the first flight with private ticket holders. Over the following year Unity flew roughly monthly. Its seventh and final commercial spaceflight, Galactic 07, took place on 8 June 2024. Virgin Galactic then retired Unity from spaceflight to conserve cash for its next-generation Delta-class ships, which it designed to fly more often and more cheaply. The company lost 279 million dollars in 2025 while flying no one. In August 2026 it told investors that the first Delta ship would enter service in February 2027, a few months later than it had said earlier in the year, with the second following in March. The Blue Origin path Blue Origin followed a very different path, financed by the personal fortune of Jeff Bezos rather than public markets. Founded in 2000, it worked in near-silence for a decade and a half. New Shepard first flew in 2015, and in November of that year its booster landed itself vertically, the first time a rocket that had gone to space returned for a controlled landing. The company then flew more than a dozen uncrewed test flights, repeatedly reusing boosters and testing the capsule's escape system in flight, before carrying people. The first crewed flight, on 20 July 2021, carried Jeff Bezos, his brother Mark, the 82-year-old aviation pioneer Wally Funk and an 18-year-old Dutch student, Oliver Daemen, whose father had bought the seat. A fourth seat had been auctioned for 28 million dollars, but the winner deferred. Crewed flights followed at intervals: William Shatner flew in October 2021, and a series of mixed crews of paying customers and sponsored guests flew through 2022. On 12 September 2022 an uncrewed New Shepard carrying research payloads suffered a failure of its engine nozzle shortly after launch. The capsule's escape motor fired as designed and pulled it clear, and it landed safely under parachutes; the booster was lost. The fleet was grounded for fifteen months while the company and the FAA investigated. Uncrewed flights resumed in December 2023 and crewed flights in May 2024. From then on New Shepard flew more steadily than ever, including an all-female crew in April 2025 that drew enormous media attention. Its thirty-eighth and last flight before the pause took place in January 2026. Blue Origin said it had carried 98 people above the Kármán line and flown more than 200 scientific payloads. Eight days after that flight it announced that New Shepard would be paused for at least two years so that resources could go to the company's lunar lander programme. The escape-system activation in 2022 is the mirror image of the SpaceShipTwo accident, and the contrast is illuminating. A capsule on top of a rocket can be designed with an escape system that separates the passengers from a failing booster; a winged vehicle with passengers inside the rocket stage generally cannot. That design choice does not make one vehicle safe and the other unsafe, but it changes the structure of risk and therefore the arguments each operator can make to passengers, insurers and regulators. The companies that did not make it The two survivors are the visible part of a larger history. The Tauri Group's 2012 market study for the FAA and Space Florida counted eleven suborbital reusable vehicles in development or operation by six companies. Most never flew people. XCOR Aerospace, which designed a two-seat rocket plane called Lynx and moved its headquarters to Midland, Texas, with the help of local economic development incentives, halted Lynx development in 2016 and filed for bankruptcy in 2017. Armadillo Aerospace, founded by the video-game developer John Carmack, went into what its founder called hibernation in 2013. Rocketplane, an earlier Oklahoma venture, went bankrupt in 2010. Beyond rockets, Space Perspective's balloon venture collapsed in 2025, as Chapter 1 described. These failures matter for two reasons. First, they reveal survivorship bias in how the industry is discussed. The two companies that reached commercial service were financed by two of the richest people alive, and even they took around two decades. Companies financed by conventional venture capital did not survive the development period. Second, several of the failures left behind public investments, notably in Midland and in Oklahoma, whose value depended on the operator that never came. That pattern is the subject of Chapter 7. What the history teaches about development risk Four lessons emerge from this history. The first is that the cost of developing a passenger suborbital system was underestimated by something like an order of magnitude. SpaceShipOne was built for a sum widely reported at around 25 million dollars. The passenger-carrying successor and its programme consumed Virgin Group funding reported at roughly a billion dollars before the company went public, and hundreds of millions of dollars a year afterwards. Blue Origin's spending on New Shepard has never been published, but it was sustained by Jeff Bezos's sale of Amazon stock, which he said in 2017 ran at about a billion dollars a year for the company as a whole. The difference between a prize-winning demonstrator and a certified-for-the-public service is not a matter of scale. It lies in reliability, maintainability, training, ground operations and the organisation needed to fly safely on a schedule. The second lesson is that turnaround, not first flight, is the hard part. Unity, once in service, flew roughly once a month. New Shepard, even in its best years, flew crewed missions every few weeks at most. These cadences are a long way below what their business plans need. The limiting factors were inspection, refurbishment of engines and structures, availability of carrier aircraft, weather, and the size of the specialised workforce. Virgin Galactic's decision to retire Unity in favour of an entirely new ship design, rather than build more copies of it, is an admission that the first design could not be turned around fast enough to make money. The third lesson is that events outside the flight cause the longest delays. The 2007 explosion occurred on the ground; the airspace episode of 2021 involved no injury; the 2022 New Shepard failure harmed nobody. Yet each produced months or years without service, as investigations proceeded and designs were revised. Operators and investors tend to model risk as the probability of a fatal accident. The more frequent and in aggregate more expensive risk is the grounding: a non-fatal anomaly that halts revenue while fixed costs continue. The fourth lesson is that the owners' strategic priorities can stop service as surely as an accident. Blue Origin's 2026 pause was not caused by any problem with New Shepard. It was caused by the fact that its owner decided the company's scarce engineers were more valuable building a lunar lander for NASA. For a customer holding a reservation, or a community that has invested in a spaceport, the effect of a strategic pause is identical to that of a grounding. It is a risk that no conventional safety analysis captures, and it exists because in each case passenger suborbital flight is a small part of a larger enterprise controlled by a single individual. The view from the waiting list The history looks different from the position of a customer. Someone who paid a deposit to Virgin Galactic in 2005, at the age of fifty, would have been sixty-eight by the time private customers first flew in 2023, and if still waiting in 2026 would be in their early seventies. Over so long a wait, some customers asked for their money back, reportedly including a number in the months after the 2014 accident. Those who stayed became, in effect, long-term unsecured creditors of a development programme, holding a claim to a future service that had no fixed date. The operators worked hard to make the wait tolerable, through events, facility tours, updates from engineers and the cultivation of a community of future astronauts. That effort was commercially rational, because a customer who remains engaged is less likely to demand a refund and more likely to recommend the product. But it also illustrates something important about the business. For most of its history, the product that Virgin Galactic actually delivered to most of its customers was not a spaceflight but membership of a club waiting for one. The value of that membership depended entirely on confidence that the flight would eventually come, which is why each accident, delay and announcement moved not only the share price but the mood of the waiting list. An objection: was the regulator to blame? A common response to this history, particularly within the industry, is that the delays were imposed from outside: that licensing, environmental review and investigation processes slowed companies down, and that a lighter regulatory touch would have brought passengers to space sooner. The objection deserves a fair hearing, because it bears directly on the policy questions of Chapters 6 and 7. There is some truth in it. Each licence modification, airspace arrangement and environmental assessment takes time, and in periods when the FAA's commercial space office was stretched thin, operators waited for approvals. The August 2025 executive order discussed in Chapter 6 was a response to exactly such complaints, though these came mainly from the orbital launch industry. Mishap investigations, even when led by the operator, cannot conclude until the regulator accepts the findings. But the weight of the evidence runs the other way. The FAA was, by statute, largely barred from regulating passenger safety throughout the whole period, and its licensing focused on public safety, where suborbital vehicles flying over remote deserts pose comparatively modest risks. The longest delays in Virgin Galactic's history followed a ground explosion, an in-flight breakup and a decision to rebuild its fleet, all of which were engineering and business events. Blue Origin's fifteen-month grounding after 2022 followed a failure of its own engine nozzle, and its 2026 pause followed a decision of its own owner. The NTSB's finding after the 2014 accident was, if anything, that the regulator had been too permissive, granting waivers of its own hazard-analysis requirements without adequate scrutiny. The industry's timelines were set by the difficulty of building and operating a passenger rocket, not by the burden of oversight. That matters because it suggests that loosening regulation further would do little to raise flight rates, while it might weaken the credibility on which demand, as the next chapter shows, partly depends. Valuing a spaceflight company in the light of its history For an investor or a public body, the history suggests a discipline. Any passenger spaceflight plan should be evaluated not by its projected steady state, but by the probability-weighted path it must traverse to get there: development cost overruns, groundings of a year or more, redesigns that retire an entire vehicle generation, and strategic pauses. The evidence of two decades is that each of these has occurred at least once to every company that reached service. Virgin Galactic's own planning for Delta implicitly acknowledges this. The company has emphasised that the new ships are designed for much faster turnaround and a service life of hundreds of flights, precisely because Unity's economics were never viable. None of this means the business cannot work. It means that the price of admission to the industry has been paid in time and capital on a scale that makes the ticket price, however high it looks, a small part of the story. The next two chapters turn to the question of how that ticket price is set, and what the demand behind it can bear. Chapter 3: Price, Demand and the Elasticity Puzzle In 2005 a seat on Virgin Galactic's future spaceship cost 200,000 dollars. In 2026 the company sold a new tranche of fifty seats at 750,000 dollars each and reported that it sold out ahead of schedule. Over two decades the price nearly quadrupled, far faster than inflation, while the product was repeatedly delayed. The company has announced that each future tranche will be priced higher than the one before. On the face of it this is a textbook case of inelastic demand: raise the price, and people keep buying. The face of it is misleading. What looks like demand that does not respond to price is better understood as a queue being rationed by price, in a market where supply has been close to zero for most of its history. The genuine price elasticity of demand for suborbital flight, the thing that will determine whether the business can grow beyond a few hundred customers a year, has barely been tested. This chapter explains why, examines what the major market studies have and have not established, and argues that the most important pricing decisions in this industry still lie ahead. A short history of the price Table 2 sets out Virgin Galactic's published seat prices over time. The company is the only operator that has priced openly and repeatedly, which makes its record the best available natural experiment in suborbital pricing. Table 2. Virgin Galactic published seat prices, 2005 to 2026. Period Price per seat Context Service status at the time 2005 to 2013 200,000 dollars Initial sales to early "founder" customers Vehicle in development 2013 to 2014 250,000 dollars Price raised during powered flight test Vehicle in test; sales paused after the 2014 accident August 2021 450,000 dollars Sales reopened after Branson's flight Unity in test; commercial service not yet begun 2023 600,000 dollars Price for later reservations Unity entering commercial service 2026 750,000 dollars Fifty-seat tranche after a two-year sales pause No flights since June 2024; Delta ships in production Sources: Virgin Galactic announcements and shareholder letters; The Register (April 2026); Space.com (August 2026). Two features of this record stand out. First, the largest price rises came at moments of high publicity and low supply: immediately after Branson's own flight in 2021, and in 2026, when the company had not flown for nearly two years and would not fly again for many months. Second, the company paused sales for long periods, which means it was not trying to discover the market-clearing price at all. It was rationing a scarce future inventory among a pool of interested buyers, raising the price each time it offered a new block. The customer base built up under this approach is a stack of cohorts who paid different prices. In 2026 Virgin Galactic reported about 675 people holding reservations, and a backlog of future spaceflight revenue of more than 240 million dollars. Dividing the one by the other is only a rough guide, since the company's backlog accounting need not map exactly onto individual reservations, but it suggests an average contracted price in the region of 350,000 dollars, reflecting the large number of early customers who bought at 200,000 or 250,000 dollars. The company has told investors that it expects an average price of about 600,000 dollars per seat across the Delta fleet's early service. These figures show the double edge of rising prices: every new tranche lifts the average, but the old reservations must still be flown at the old prices, using capacity that could otherwise be sold at the new ones. What price elasticity would mean here Price elasticity of demand is the percentage change in the quantity demanded divided by the percentage change in price. If a ten per cent price cut raises sales by twenty per cent, demand is elastic, with an elasticity of about minus two, and revenue rises when prices fall. If the same cut raises sales by only five per cent, demand is inelastic and revenue falls. For a business with high fixed costs and low marginal costs, like an airline or a spaceship operator, the shape of this relationship across a wide range of prices determines the optimal strategy. The difficulty in suborbital tourism is that the demand curve has two quite different regions. At the top is a thin layer of very wealthy buyers for whom the price of a ticket is a small fraction of their annual income, and for whom the scarcity and prestige of the experience are part of its value. Demand in this layer is close to perfectly inelastic across the prices so far charged, and may even show a "Veblen" pattern in which a higher price increases the good's appeal as a status marker. Below it lies a much larger population of affluent people, dollar millionaires and the upper professional classes, who would pay a large but not unlimited sum for a once-in-a-lifetime experience, and for whom price is decisive. The first layer is finite. The property consultancy Knight Frank estimated in its 2025 Wealth Report that almost 630,000 people worldwide had net assets of at least 30 million dollars. Not all of them want to go to space, many are too old or unwell to do so, and many of those who want to have already signed up. Surveys conducted over the past two decades suggest that a meaningful but minority share of the very wealthy expresses serious interest, and a smaller share will ever put down a deposit. The existing reservation lists, a few hundred at Virgin Galactic and an undisclosed number at Blue Origin, may be a substantial fraction of the high-price market that exists today. When those lists are exhausted, the business will depend on the second layer, where elasticity is high and where current prices would sell very few seats. The market studies, and what they got right The most influential early attempt to quantify demand was the Futron Corporation's 2002 study, based on a survey conducted with Zogby International of 450 wealthy Americans. Futron asked respondents about their interest at various price points and built a forecast in which suborbital seats would start at 100,000 dollars and fall to 50,000 dollars by 2021. On those assumptions the study projected that more than 15,000 passengers a year would fly by 2021, generating revenue of about 785 million dollars. A 2006 revision, allowing for a later start and higher initial prices of 200,000 dollars, cut the 2021 projection to just over 13,000 passengers and 676 million dollars. Relaxing the fitness assumptions to include people of average fitness under sixty-five roughly doubled the projection. In 2012 the Tauri Group produced a ten-year demand forecast for suborbital reusable vehicles for the FAA and Space Florida. It covered not only tourism but research, education, media and other uses, and offered three scenarios: a constrained case worth about 300 million dollars over the decade, a baseline case worth more than 600 million dollars with daily flights, and a growth case worth about 1.6 billion dollars. Tourism was the largest single component in every scenario. Investment banks followed with larger numbers. In 2019 UBS estimated that space tourism could be worth about three billion dollars a year by 2030, and in 2021 it raised the figure to about four billion. Commercial market-research firms have produced a proliferation of forecasts since, most of which are neither transparent in method nor independent of one another. Measured against what has happened, the headline forecasts were wrong by orders of magnitude on timing and volume. By the end of 2021 fewer than two dozen people had flown on commercial suborbital vehicles, most of them in the second half of that year and several of them company founders, employees or guests rather than paying customers; the forecast had been for thousands a year. Even by 2026 the cumulative total across both American operators was well under two hundred people. But it is important to be precise about why the forecasts failed. They were not refuted by a shortage of interested buyers. They were refuted by the absence of supply and by prices that never fell. Futron's projections rested explicitly on prices dropping to 50,000 dollars, which is a fifteenth of the price actually charged in 2026. The forecasts described a demand curve; the industry has so far operated only at its very top. Academic work has made the same point more rigorously. A team led by Geoffrey Crouch, in a 2009 paper in the journal Tourism Management, used a discrete choice experiment, in which respondents choose between hypothetical flights with systematically varied attributes, to model how consumers trade price against features such as duration, altitude, training and safety record. Such studies consistently find that price and perceived safety are among the strongest determinants of choice, and that the potential market expands sharply as price falls. A 2019 analysis by Markus Guerster, Edward Crawley and Richard de Neufville at MIT went further, simulating thousands of possible demand futures and asking which strategy for vehicle size and fleet expansion would perform best. Their conclusion was that suborbital tourism could probably be made commercially attractive, but only by managing development flexibly in the face of deep uncertainty, and that strategies built around larger vehicles and aggressive expansion had higher expected value but also a higher chance of losing money. Why operators keep raising prices If demand expands so sharply at lower prices, why has Virgin Galactic kept raising them? The answer lies in the economics described in the next chapter, but the pricing logic can be stated simply. While capacity is severely constrained, lowering the price does not increase the number of flights; it only reduces revenue per seat and lengthens the queue. The profit-maximising strategy for a supply-constrained monopolist is to sell each scarce seat to the buyer willing to pay most for it, and to use the resulting revenue and publicity to fund expansion. That is precisely what the company has done, and it is rational. It also serves a financial purpose. For a publicly listed company burning cash, a rising price per seat is a visible signal of demand that supports the share price and the ability to raise equity. Virgin Galactic raised about 134 million dollars through share sales in the second quarter of 2026 alone. Announcing a sold-out tranche at a record price is part of the case it makes to the capital markets. There is a less comfortable reading too. Deposits paid for seats that will not fly for years are, in effect, interest-free loans from customers. Virgin Galactic's reservation terms have historically required a large deposit, only part of it refundable in some circumstances. The longer the delay between sale and flight, the more the customer is bearing development risk without being compensated for it, and the more the price increases represent the operator's cost of capital being paid by its most enthusiastic supporters. Novelty, reputation and the durability of demand A demand curve is not fixed. Two forces could shift the demand for suborbital flight downward over time, and both deserve more attention than the market studies gave them. The first is the decay of novelty. Part of what the early customers bought was the chance to be among the first few hundred private people to go to space. That value falls as the number of people who have flown rises. The FAA's decision to stop awarding commercial astronaut wings, described in Chapter 1, was a small institutional recognition of the same effect. Demand that rests on scarcity is, by definition, eroded by supply. Operators will increasingly have to sell the experience on its intrinsic merits, the view of the Earth, the sensation of weightlessness, the shared adventure, rather than on its exclusivity. There is good reason to think those merits are substantial; almost every person who has flown describes the view as profound. But the price that intrinsic value will command is unknown, and may well be lower than the price exclusivity has commanded. The second is reputation. Space tourism has attracted criticism as a conspicuous indulgence of the very rich, and that criticism rises with the visibility of each flight. New Shepard's all-female flight of April 2025, which carried the singer Katy Perry and the broadcaster Gayle King among others, drew enormous attention and a substantial public backlash that treated it as a celebrity stunt. Environmental critiques add weight. A 2022 study in the journal Earth's Future by Robert Ryan, Eloise Marais and colleagues estimated that black carbon, or soot, emitted by rockets directly into the upper atmosphere is roughly 500 times more effective at warming the climate than the same soot released at the surface or by aircraft, and specifically flagged the growth of space tourism as a concern. Suborbital flights are too few for their total emissions to be large today, but the argument is about trajectory, and it is heard by precisely the affluent, educated customers the industry needs to reach. For some buyers a spaceflight is a status good; for others, if public opinion shifts, it may become a mild embarrassment. Neither force is likely to empty the reservation lists soon. But both suggest that treating today's willingness to pay as a stable foundation for a decade of pricing would be a mistake. The market has been sold on scarcity and spectacle. Its long-run size depends on whether it can be sold on something more lasting. When price discovery will actually begin The real test of the market will come only when supply stops being the binding constraint. On Virgin Galactic's own projections, two Delta ships flying at a combined rate of ten or more flights a month with six passenger seats each would offer about 720 seats a year, which is roughly the size of the entire current reservation list. Within a year or so of full service, if the ships perform as planned, the company would exhaust its queue and have to find several hundred new customers every year. At that point the pricing question changes character. The company can keep prices high and fly fewer passengers, filling seats with research payloads and institutional customers where it can; it can lower prices to reach the deeper and more elastic layer of demand; or it can segment the market, as airlines and luxury hospitality do, by offering different experiences at different prices. Each choice has consequences. Lower prices would require flights to be much cheaper to produce, which in turn requires higher cadence. Segmentation requires distinguishable products, which a fixed six-seat cabin makes difficult. Holding prices high while the fleet sits idle would defeat the purpose of building ships designed for five hundred flights each. Competitors will shape the answer. Deep Blue Aerospace in China has already sold seats at about 210,000 dollars for a service it hopes to begin in 2027, less than a third of Virgin Galactic's current price. Blue Origin, when New Shepard returns, will be able to set its own price from a position of having flown nearly a hundred people. Virgin Galactic's chief executive has said that he believes Blue Origin charged between one and two million dollars a seat, but Blue Origin has never confirmed its prices. A market with three or four suppliers flying regularly would look very different from the near-monopoly queues of the past. The lesson for pricing strategy Suborbital pricing has so far been a problem of rationing, not of discovery. The elasticity that matters, the responsiveness of the affluent middle of the wealth distribution to a lower price, remains largely unmeasured except in surveys and choice experiments, which are notoriously poor at predicting how people actually spend very large sums. What the surveys agree on is that the market below the current price is much larger than the market at it, and that perceived safety matters as much as price in determining who will fly. That last finding connects demand directly to liability and regulation. A buyer's willingness to pay depends on their belief about the risk, and that belief depends on the record of flights, the regime of disclosure and the credibility of whoever vouches for the vehicle. The price an operator can charge, and the size of the market it can reach, are therefore not only matters of marketing. They are downstream of the law, which is why the middle chapters of this booklet turn to it. First, though, it is necessary to see why the cost side of the business pushes so hard towards higher volume. Hashtags: #TheBusinessOfSpaceTourism #SpaceTourism #PassengerSpaceflight #CommercialHumanSpaceflight #SuborbitalTourism #OrbitalTourism #SpaceflightDemand #TicketPricing #PriceElasticity #FlightCadence #UnitEconomics #DevelopmentRisk #SpaceflightLiability #InformedConsent #ReciprocalWaivers #ThirdPartyLiability #OccupantSafety #FAALearningPeriod #SpaceflightRegulation #SpaceportInfrastructure #PublicSpaceportFinance #PrivateAstronautMissions #CommercialSpaceStations #SpaceflightMarketDemand #FutureOfSpaceTourism
- The Cancer-Diet Connection (Unpacking Nutritional Oncology)
Download the Book (PDF): Introduction Every student who opens a textbook of nutritional oncology meets the same problem within the first few weeks. One chapter asks you to understand why a sausage cooked over an open flame might, over thirty years, nudge the probability that a colonic crypt cell acquires a mutation in the APC gene. The next asks you to calculate the protein requirement of a woman who has lost eleven kilograms during chemotherapy for pancreatic cancer and who now cannot face more than a few spoonfuls of soup. Both sit under the same heading, "diet and cancer", and both are examinable. Yet they involve different questions, different kinds of evidence, different time scales and, crucially, different kinds of advice. Students who try to hold them in a single mental compartment tend to make predictable errors: they tell a cachectic patient to cut out red meat and sugar, or they describe processed meat as "as dangerous as smoking" because both appear in the same carcinogen category. This companion is built on one controlling idea. Nutrition in oncology does two fundamentally different jobs on either side of a diagnosis. Before diagnosis, diet shifts the probability that cancer develops across a population over decades, and the evidence is epidemiological and mechanistic. After diagnosis, nutrition protects an individual patient's capacity to tolerate, complete and recover from treatment, and the evidence is clinical. Most errors in the field, and most of its myths, come from carrying the logic of one pathway into the other. Keep the two pathways distinct, understand the biology that connects them, and the subject becomes coherent rather than overwhelming. Why two pathways, and why they share a foundation The first pathway is prevention. Its central question is: which dietary exposures alter the risk of which cancers, by how much, and through what mechanism? The answers come from cohort studies of hundreds of thousands of people, from pooled analyses, from occasional large trials, from laboratory work on DNA adducts and hormone signalling, and from expert panels that grade all of this. The unit of concern is the population and the lifetime. A change in relative risk of fifteen or twenty per cent is important at this scale because it applies to millions of people, even though for any single person the absolute change may be small. The second pathway is treatment. Its central question is: how do we keep this patient nourished, strong and functional enough to receive the surgery, chemotherapy, radiotherapy or immunotherapy that offers them the best outcome, and how do we manage the symptoms that get in the way? The evidence comes from clinical trials, consensus definitions and practice guidelines, particularly those of the European Society for Clinical Nutrition and Metabolism (ESPEN). The unit of concern is the individual and the next few weeks. Here the dietitian's priorities change. A patient who is losing weight rapidly needs energy and protein, and the prevention message of "limit red meat and sugary foods" can actively harm them if it is applied without thought. The two pathways are not unrelated. They share a biological foundation. The same processes that explain how diet can promote cancer, such as chronic inflammation, insulin and insulin-like growth factor signalling, oxidative DNA damage, the evasion of apoptosis and the recruitment of new blood vessels through angiogenesis, also explain why tumours cause the metabolic chaos of cachexia and why certain nutritional claims made to patients are biologically implausible. A student who understands the hallmarks of cancer can reason about a new claim instead of memorising a list of approved and forbidden foods. How this guide is organised The chapters follow the two pathways in order, with a shared foundation at the start. Chapter 1 sets out the biology: the hallmarks of cancer, the stages of initiation, promotion and progression, and rigorous definitions of apoptosis and angiogenesis. These are the mechanisms you will be asked to explain in essays, and every later chapter draws on them. Chapter 2 explains how evidence about diet and cancer is generated and graded. It covers the main study designs and their weaknesses, the grading system used by the World Cancer Research Fund and American Institute for Cancer Research (WCRF/AICR) in their Continuous Update Project, and the hazard classification of the International Agency for Research on Cancer (IARC). Understanding the difference between hazard and risk will save you from one of the most common errors in student writing. Chapters 3 to 5 form the prevention pathway. Chapter 3 deals with dietary carcinogens and DNA damage: N-nitroso compounds and haem iron, heterocyclic amines and polycyclic aromatic hydrocarbons, aflatoxin, and alcohol with its metabolite acetaldehyde. Chapter 4 turns to the promotional environment created by excess body fat, including insulin, insulin-like growth factor 1, adipokines, oestrogen and chronic inflammation. Chapter 5 examines protective dietary patterns, fibre, whole grains, calcium and the phytochemicals, and explains why the supplement trials so often disappointed, before setting out the 2018 WCRF/AICR recommendations. Chapters 6 to 9 form the treatment pathway. Chapter 6 defines cancer cachexia using the 2011 international consensus, explains its mechanisms and shows how malnutrition is screened for and diagnosed. Chapter 7 translates this into requirements and nutrition support, with worked calculations and the logic of oral, enteral and parenteral routes. Chapter 8 covers the nutrition impact symptoms of chemotherapy, radiotherapy and newer systemic treatments, site by site. Chapter 9 addresses surgery: enhanced recovery, prehabilitation, immunonutrition and the specific nutritional consequences of major cancer resections. Chapter 10 brings the pathways back together in survivorship, where prevention logic partially returns, and then evaluates the popular claims that patients bring to clinic: the alkaline diet, the idea that sugar feeds cancer, fasting and ketogenic diets, high-dose antioxidants and others. The conclusion argues what follows from the whole for how you should practise and write. How to use this guide alongside your textbook The textbook of nutritional oncology you are studying covers the field in depth, drawing on epidemiology, molecular biology and clinical practice. This guide is an independent companion. It does not summarise the textbook chapter by chapter; instead it gives you a framework to hang that detail on. When you read a dense section of the textbook on, say, the metabolism of heterocyclic amines, return to Chapter 3 here to place it in the prevention pathway. When you read about parenteral nutrition in advanced disease, use Chapter 7 to see where it fits in clinical decision-making. Each chapter includes definitions you can use in exams, mechanisms written as step-by-step sequences, worked examples with hypothetical patients and notes on common pitfalls. There are six tables in the book, reserved for material that is genuinely comparative. Numbers such as guideline targets and diagnostic thresholds are given where they are well established; laboratory reference ranges vary between laboratories, so where they appear they are marked as typical. A note on tone and honesty Cancer is frightening, and nutrition is one of the few things patients feel they can control. That makes the field fertile ground for exaggeration in both directions: foods marketed as cures, and ordinary foods described as poison. A good nutrition professional holds a calm position between these. You should be able to tell a patient, honestly, that processed meat is a well-established cause of colorectal cancer and that eating it regularly does raise risk, while also telling them that the bacon sandwich they ate last week did not cause their tumour. You should be able to tell another patient, equally honestly, that sugar does not "feed" their cancer in the way websites claim, while explaining why their oncologist is watching their blood glucose during steroid treatment. That combination of precision and proportion is what this book aims to teach. Chapter 1: The Biology Beneath Both Pathways A cancer is a population of cells that has escaped the controls which normally keep cell number, position and behaviour in balance with the needs of the whole organism. That sentence is simple, but every word matters for nutrition. "Population" reminds us that a tumour evolves by natural selection among its own cells. "Escaped" reminds us that controls exist and must be broken, usually one at a time, through mutation and epigenetic change. "Balance with the needs of the organism" reminds us that cancer is a systemic disease: the tumour communicates with the immune system, the blood vessels, adipose tissue, the liver and the brain, and those conversations are exactly where diet exerts its influence before diagnosis and where malnutrition arises after it. This chapter builds the foundation you need for both pathways. It sets out the hallmarks of cancer, the classical stages of carcinogenesis, and then defines two processes with the precision examiners expect: apoptosis and angiogenesis. The hallmarks of cancer In 2000 Douglas Hanahan and Robert Weinberg proposed that the enormous diversity of cancers could be understood through a small number of acquired capabilities that almost all tumours need. Their framework was updated in 2011 and again by Hanahan in 2022. It is the standard organising scheme in oncology teaching, and it is also the best tool for asking "where could diet act?" The original six capabilities are: Sustaining proliferative signalling. Normal cells divide only when instructed by growth factors. Cancer cells supply their own growth signals, overexpress receptors, or carry mutations that lock signalling pathways in the "on" position. Mutations in KRAS and activation of the phosphoinositide 3-kinase (PI3K), AKT and mTOR pathway are classic examples. This hallmark is the one most directly touched by nutrition through insulin and insulin-like growth factor 1 (IGF-1), which signal through PI3K-AKT-mTOR. Evading growth suppressors. Tumour suppressor proteins such as the retinoblastoma protein (RB) and p53 act as brakes on the cell cycle. Loss of these brakes, through mutation or deletion, allows division to continue when it should stop. Resisting cell death. Cells with damaged DNA or abnormal signalling are normally eliminated by apoptosis. Cancer cells disable this programme. Enabling replicative immortality. Normal cells can divide only a limited number of times, partly because their telomeres shorten with each division. Cancer cells typically reactivate telomerase, allowing indefinite division. Inducing angiogenesis. A tumour beyond a very small size cannot survive on diffusion alone and must recruit new blood vessels. Activating invasion and metastasis. Cancer cells lose adhesion to their neighbours, degrade surrounding matrix, enter the circulation and colonise distant organs. The 2011 revision added two further hallmarks, deregulating cellular energetics (the metabolic reprogramming that includes the Warburg effect, discussed in Chapter 10) and avoiding immune destruction, together with two "enabling characteristics": genome instability and mutation, which generates the variation on which selection acts, and tumour-promoting inflammation, which supplies growth factors, survival signals and mutagenic reactive oxygen species. The 2022 update proposed further emerging features, including unlocking phenotypic plasticity, non-mutational epigenetic reprogramming, polymorphic microbiomes and senescent cells. For a nutrition student, the power of this framework is that it lets you map dietary exposures onto specific capabilities. Alcohol-derived acetaldehyde and meat-derived N-nitroso compounds act on genome instability. Obesity-related hyperinsulinaemia acts on proliferative signalling. Chronic inflammation from visceral adiposity or Helicobacter pylori infection acts through tumour-promoting inflammation. The gut microbiome, which diet shapes powerfully, appears in the newest list. When an essay asks you to "discuss the mechanisms by which diet influences carcinogenesis", structuring the answer around hallmarks immediately signals command of the subject. A caution: the hallmarks describe what tumours have acquired, not what caused them. A phytochemical that inhibits angiogenesis in a cell culture dish has touched a hallmark, but that does not mean eating the food that contains it will prevent cancer. We return to this gap between mechanism and outcome repeatedly. The stages of carcinogenesis The classical model of chemical carcinogenesis, developed largely from experiments on mouse skin in the mid-twentieth century, divides the process into three stages. It is simplified, but it remains useful because it distinguishes different kinds of dietary influence. Initiation is the induction of a permanent, heritable change in a cell's DNA. It usually results from a carcinogen, or more often its reactive metabolite, binding covalently to DNA to form an adduct. If the adduct is not repaired before the cell divides, DNA polymerase may insert the wrong base opposite it, and the error becomes fixed as a mutation. Initiation is rapid, irreversible and in principle can follow a single exposure. An initiated cell looks normal and may remain dormant for life. Many dietary carcinogens are procarcinogens: they are chemically inert until the body activates them. Phase I enzymes, especially the cytochrome P450 family, add or expose reactive groups; phase II enzymes such as glutathione S-transferases, UDP-glucuronosyltransferases, sulfotransferases and N-acetyltransferases then conjugate these groups to make them water-soluble for excretion. The balance matters. Some phase II reactions detoxify, but others, such as the O-acetylation of hydroxylated heterocyclic amines, generate the ultimate carcinogen. Genetic variation in these enzymes partly explains why two people with the same diet carry different risk. The key point for examinations is that initiation depends on dose, metabolic activation, DNA repair capacity and cell division. Promotion is the selective clonal expansion of initiated cells. Promoters are typically not mutagenic themselves. They act by stimulating proliferation, inhibiting apoptosis or creating a tissue environment that favours the initiated clone. Promotion is slow, requires repeated or sustained exposure, shows a threshold, and is at least partly reversible in its early phases. This is the stage at which much dietary influence probably operates. Hyperinsulinaemia, elevated free IGF-1, oestrogen excess, chronic inflammation and bile-acid-driven irritation of the colonic mucosa are all promotional forces. The reversibility of promotion is the biological basis for optimism about dietary change in adult life: reducing a promotional pressure may slow or halt the expansion of initiated clones even if the initiating mutations cannot be undone. Progression is the transition from a benign, expanding clone to a malignant tumour with the capacity for invasion and metastasis. It is driven by accumulating genetic instability, including chromosomal rearrangements and loss of further tumour suppressors. Once progression has occurred, dietary modification has much less power over the tumour's behaviour, which is one reason the evidence for diet in treating established cancer is so much weaker than for prevention. The best worked example of this multistage process is the colorectal adenoma-carcinoma sequence described by Bert Vogelstein and colleagues. Loss of the APC tumour suppressor gene allows a normal crypt to become hyperproliferative and form a small adenoma. Activating mutations in KRAS support growth to a larger adenoma. Loss of further suppressor functions, historically associated with chromosome 18q and genes such as SMAD4, advances the lesion, and loss of TP53 is associated with the transition to invasive carcinoma. The sequence usually takes a decade or more. This long window explains why the colon is both the organ where diet has its strongest established effects and the organ where screening by colonoscopy can remove precursors before they become malignant. Pitfall to avoid: do not write that "processed meat initiates colorectal cancer" as if a single mechanism were established. The honest statement is that processed meat increases colorectal cancer risk, and plausible mechanisms include genotoxic N-nitroso compounds (initiation) and haem-driven mucosal cytotoxicity and hyperproliferation (promotion). The staging model is a framework for thinking, not a claim that each exposure acts at only one stage. Apoptosis defined Apoptosis is a genetically regulated, energy-dependent programme of cell death in which a cell dismantles itself in an orderly way, without releasing its contents or provoking inflammation, and is then engulfed by neighbouring cells or macrophages. The morphological features are cell shrinkage, chromatin condensation, fragmentation of the nucleus, blebbing of the plasma membrane and packaging of the cell's contents into membrane-bound apoptotic bodies. A key molecular signal is the flipping of phosphatidylserine from the inner to the outer leaflet of the plasma membrane, which marks the cell for engulfment. Contrast this with necrosis, the uncontrolled death that follows severe injury, in which cells swell and rupture, spill their contents and trigger inflammation. The distinction matters because apoptosis is a tumour-suppressive process, while necrosis within tumours can fuel inflammation. Apoptosis is executed by caspases, cysteine proteases that cleave their targets after aspartate residues. They exist as inactive precursors and are activated in cascades. Two main pathways converge on the executioner caspases. The intrinsic (mitochondrial) pathway responds to internal stress, including DNA damage, oxidative stress, growth factor withdrawal and oncogene activation. Its gatekeepers are the BCL-2 family of proteins. Anti-apoptotic members such as BCL-2 and BCL-xL hold pro-apoptotic effectors BAX and BAK in check. Stress signals activate "BH3-only" proteins such as PUMA, NOXA and BIM, which tip the balance. BAX and BAK then permeabilise the outer mitochondrial membrane, releasing cytochrome c into the cytosol. Cytochrome c binds APAF-1 to form a structure called the apoptosome, which activates the initiator caspase-9. Caspase-9 activates the executioner caspases-3 and -7, which cleave hundreds of cellular proteins and bring about the characteristic changes. The extrinsic (death receptor) pathway responds to signals from outside the cell. Ligands such as Fas ligand, tumour necrosis factor alpha (TNF-α) and TRAIL bind their death receptors on the cell surface. The receptors recruit adaptor proteins such as FADD, forming a death-inducing signalling complex that activates the initiator caspase-8. Caspase-8 directly activates caspase-3, and can also cleave the BH3-only protein BID, linking the extrinsic pathway into the intrinsic one to amplify the signal. Cytotoxic T cells and natural killer cells use this route to kill tumour cells. The protein p53 sits at the centre of apoptotic control, which is why it is called "the guardian of the genome". When DNA is damaged, p53 is stabilised and acts as a transcription factor. It can halt the cell cycle, mainly through p21, allowing time for repair; if damage is irreparable, it induces pro-apoptotic genes including PUMA, NOXA and BAX. Loss of p53 function, which occurs in roughly half of human cancers, allows cells with damaged DNA to survive and divide. The link to diet is direct in at least one case: aflatoxin B1, discussed in Chapter 3, produces a characteristic mutation at codon 249 of TP53 in liver cancers from regions of high exposure. Cancer cells evade apoptosis in several ways: overexpressing BCL-2 (as in follicular lymphoma, where a chromosomal translocation drives it), losing p53, overexpressing inhibitor of apoptosis proteins, or downregulating death receptors. Survival signalling through PI3K-AKT also suppresses apoptosis; AKT phosphorylates and inactivates the pro-apoptotic protein BAD. This is one route by which insulin and IGF-1 connect nutritional state to cell survival. Apoptosis matters in the treatment pathway too. Many chemotherapy agents and radiotherapy work largely by damaging DNA and triggering apoptosis in dividing cells. Rapidly dividing normal tissues, including the gut epithelium, oral mucosa and bone marrow, undergo apoptosis as well, which is the biological basis of mucositis, diarrhoea and neutropenia in Chapter 8. This also explains a concern discussed in Chapter 10: if antioxidant supplements in high doses reduce the oxidative damage that radiotherapy relies upon, they could in principle reduce its effectiveness. Angiogenesis defined Angiogenesis is the formation of new blood vessels from pre-existing vessels, by sprouting and branching of capillaries. It should be distinguished from vasculogenesis, the formation of vessels de novo from endothelial precursor cells, which dominates in embryonic development. In healthy adults angiogenesis is largely quiescent, switched on transiently in wound healing, the menstrual cycle and pregnancy. Why does a tumour need it? Oxygen and nutrients can diffuse only a short distance through tissue, generally cited as around 100 to 200 micrometres from a capillary. A cluster of tumour cells can therefore grow only to a very small size, roughly one to two millimetres, before its central cells become hypoxic. Such microscopic tumours may persist for years in a dormant state, with proliferation balanced by cell death. The transition to vascularised, expanding growth is called the angiogenic switch. It occurs when the balance between pro-angiogenic and anti-angiogenic factors tips in favour of the former. The central molecular sequence runs as follows: As tumour cells outgrow their blood supply, oxygen tension falls. Under normal oxygen levels, the transcription factor subunit hypoxia-inducible factor 1 alpha (HIF-1α) is continually hydroxylated by prolyl hydroxylase enzymes, which require oxygen. Hydroxylated HIF-1α is recognised by the von Hippel-Lindau (VHL) protein and targeted for proteasomal degradation. Under hypoxia, hydroxylation fails, HIF-1α is stabilised, moves to the nucleus and pairs with HIF-1β. The HIF complex switches on genes including vascular endothelial growth factor A (VEGF-A), as well as genes for glycolytic enzymes and glucose transporters. The latter link hypoxia to the metabolic reprogramming of tumours. VEGF-A diffuses to nearby vessels and binds its receptor, principally VEGFR-2, on endothelial cells. Endothelial cells loosen their junctions, the basement membrane is degraded by matrix metalloproteinases, and a leading "tip cell" migrates towards the VEGF gradient, followed by proliferating "stalk cells" that form a new lumen. The new vessels are recruited into circulation. Oncogene activation and loss of tumour suppressors can also drive VEGF production independently of hypoxia, and loss of VHL in clear cell renal carcinoma produces constitutive HIF activity, which explains why those tumours are so highly vascular. Endogenous inhibitors, such as thrombospondin-1, oppose the process; p53 promotes thrombospondin-1 expression, another link between tumour suppressor loss and the angiogenic switch. Tumour vessels are abnormal: tortuous, leaky, poorly covered by supporting pericytes and irregular in flow. This produces patchy hypoxia and high interstitial pressure within tumours, which can impair the delivery of chemotherapy and the efficacy of radiotherapy, since radiation-induced DNA damage depends partly on oxygen. Anti-angiogenic drugs such as bevacizumab, an antibody against VEGF-A, and tyrosine kinase inhibitors against VEGF receptors are now part of standard treatment for several cancers. Their side effects, including hypertension, proteinuria and impaired wound healing, have nutritional relevance, particularly around surgery. Where does diet come in? Obesity is pro-angiogenic: expanding adipose tissue itself requires new vessels, and adipose tissue and its resident macrophages secrete VEGF, leptin and inflammatory cytokines. Many phytochemicals, including compounds from green tea, soy and cruciferous vegetables, show anti-angiogenic effects in cell and animal models. As with apoptosis, the laboratory signal is real but the translation to human outcomes through normal diets is unproven. In an exam, credit comes from stating both the mechanism and the evidential limit. Putting the foundation to work The value of this chapter is that it gives you a set of questions to ask of any dietary claim or clinical problem. Is the proposed effect on initiation, promotion or progression? Which hallmark does it touch? Does it act through DNA damage, proliferative signalling, apoptosis, angiogenesis, inflammation or immunity? Is the evidence mechanistic only, or has it been tested in people with cancer outcomes? The next chapter supplies the second half of that toolkit: how to judge the strength of evidence that links diet to cancer in human populations. Chapter 2: Reading the Evidence on Diet and Cancer The prevention pathway rests on a body of evidence that is large, uneven and often misreported. Newspaper headlines swing between "coffee causes cancer" and "coffee prevents cancer" within the same year. Students need a way to judge what is known with confidence and what is merely suggested. This chapter explains how diet-cancer evidence is produced, why it is so difficult to produce well, and how two influential international bodies grade it: the World Cancer Research Fund and American Institute for Cancer Research (WCRF/AICR), and the International Agency for Research on Cancer (IARC). The central lesson is that these two bodies answer different questions, and confusing them is one of the most common errors in the field. Why diet and cancer is a hard question Several features of the problem make it unusually difficult. Long latency. Most solid cancers develop over decades. The dietary exposure that matters may have occurred twenty or thirty years before diagnosis, perhaps even in childhood or adolescence. A study that measures diet for five years before diagnosis may miss the relevant window. Measurement error. Diet is a complex, changing, poorly remembered exposure. Food frequency questionnaires, the workhorse of large cohort studies, ask people to estimate their usual intake of dozens of foods over the past year. They rank people reasonably well but measure absolute intake poorly. Random error of this kind tends to weaken associations towards the null, so real effects may be underestimated. Systematic error, such as the tendency of people with obesity to under-report energy intake, can bias results in either direction. Correlated exposures. People who eat a lot of processed meat also tend, on average, to eat fewer vegetables, drink more alcohol, smoke more and exercise less. Statistical adjustment for these confounders is imperfect, particularly when the confounder itself is measured with error. Residual confounding is therefore a constant worry. Small effect sizes. Many dietary associations are in the range of relative risks between about 0.8 and 1.3. Effects of that size are easily produced or hidden by bias. Compare this with smoking and lung cancer, where relative risks for heavy smokers are of the order of twenty or more, and there is little doubt about causation. Heterogeneity of cancer. "Cancer" is not one disease. Diet may affect oesophageal adenocarcinoma and oesophageal squamous cell carcinoma in opposite directions, or premenopausal and postmenopausal breast cancer differently. Evidence must be assessed site by site and sometimes subtype by subtype. The main study designs Ecological and migrant studies compare populations. They generated many of the original hypotheses. International variation in cancer rates is striking: colorectal and breast cancer were historically far more common in high-income Western countries, while stomach and liver cancers were more common in parts of East Asia and sub-Saharan Africa. Studies of Japanese migrants to Hawaii and the continental United States showed that, over one or two generations, stomach cancer rates fell and colorectal cancer rates rose towards those of the host population. Because genes do not change that fast, these studies established that environment, including diet, matters. But ecological associations cannot be attributed to individuals (the ecological fallacy) and are heavily confounded by everything else that differs between populations. Case-control studies compare the past diet of people with cancer with that of people without it. They are efficient for rare cancers but vulnerable to recall bias: a person newly diagnosed with cancer may search their memory differently, perhaps over-reporting foods they have read are harmful. Selection of appropriate controls is also difficult. Many early case-control studies reported protective effects of fruit and vegetables that were considerably weaker or absent in later prospective studies. Prospective cohort studies measure diet in healthy people and follow them for years to see who develops cancer. Because diet is recorded before diagnosis, recall bias is avoided. Large cohorts include the European Prospective Investigation into Cancer and Nutrition (EPIC), which recruited around half a million people across ten European countries, the NIH-AARP Diet and Health Study, the Nurses' Health Study and the Health Professionals Follow-up Study, and UK Biobank. Cohort studies are now the backbone of the evidence, but they remain observational and so vulnerable to confounding and measurement error. Pooled analyses and meta-analyses of cohorts increase statistical power and allow dose-response relationships to be estimated. Randomised controlled trials (RCTs) are the only design that deals convincingly with confounding. They are, however, very difficult for diet and cancer: people must be randomised to sustain a dietary change for many years, adherence drifts, and the number of cancer outcomes needed requires tens of thousands of participants. The Women's Health Initiative Dietary Modification Trial randomised postmenopausal women to a low-fat dietary pattern and did not demonstrate a statistically significant reduction in invasive breast cancer in its main analysis, although interpretation has been debated because the achieved difference in fat intake was smaller than planned. Trials of single nutrients as supplements are easier to run, and their results, discussed in Chapter 5, have been sobering: several showed no benefit and some showed harm. Mendelian randomisation uses genetic variants that influence an exposure, for example variants associated with higher body mass index or with lower alcohol metabolism, as natural randomising instruments. Because genes are allocated at conception, they are less susceptible to confounding by lifestyle. Mendelian randomisation studies have strengthened the causal case for body fatness in several cancers. They rely on assumptions, notably that the gene influences cancer only through the exposure, that cannot always be verified. Mechanistic studies in cell lines, animals and human biomarkers establish biological plausibility, and are essential for IARC evaluations. They cannot, on their own, show that a food at normal human intakes changes cancer risk. The WCRF/AICR Continuous Update Project The WCRF and AICR have produced major reports on diet, nutrition, physical activity and cancer since 1997, with a second report in 2007. The Continuous Update Project (CUP) replaced one-off reports with a rolling process: a team systematically searches for and meta-analyses new studies, cancer site by cancer site, and an independent expert panel judges the evidence. The Third Expert Report, Diet, Nutrition, Physical Activity and Cancer: a Global Perspective, was published in 2018 together with the Cancer Prevention Recommendations described in Chapter 5. The CUP panel grades evidence using defined criteria. Evidence is described as strong or limited. Strong evidence has three categories: • Convincing: strong enough to support a judgement of a causal relationship, and robust enough that it is unlikely to be modified by new evidence. Criteria include evidence from more than one study type, from at least two independent cohorts, no substantial unexplained heterogeneity, good-quality studies that exclude random and systematic error with confidence, a biological gradient (dose-response), and strong and plausible experimental evidence. • Probable: strong enough to support a judgement of a probably causal relationship, but with somewhat less consistency or quality than convincing. • Substantial effect on risk unlikely: strong evidence that an exposure does not have a substantial effect on risk. Limited evidence has two categories: • Limited, suggestive: generally consistent in direction but too limited in quantity or quality to judge causality. • Limited, no conclusion: too sparse, inconsistent or poor to draw any conclusion. The panel also considers whether strong evidence justifies public health recommendations. Only "convincing" and "probable" judgements generally underpin recommendations. For example, the CUP judged the evidence that processed meat increases colorectal cancer risk to be convincing, and that red meat increases it to be probable. It judged that wholegrains and foods containing dietary fibre probably protect against colorectal cancer, and that alcoholic drinks are a convincing cause of cancers of the mouth, pharynx, larynx, oesophagus (squamous cell carcinoma), colorectum and postmenopausal breast, among others. The key point for your writing is to use the grade: "the WCRF/AICR panel judged the evidence to be probable" is a much stronger sentence than "studies show". IARC and the classification of hazards IARC, part of the World Health Organization, runs the IARC Monographs programme, which identifies agents that can cause cancer in humans. Working groups of independent scientists review the published evidence from three streams: cancer in humans, cancer in experimental animals, and mechanistic evidence. Since the 2019 revision of the Monographs Preamble, mechanistic evidence is organised around the key characteristics of carcinogens, such as being electrophilic or metabolically activated, being genotoxic, inducing oxidative stress, inducing chronic inflammation, being immunosuppressive, and modulating receptor-mediated effects. As Table 1 sets out, the working group then places the agent into one of four groups. Table 1. IARC Monographs classification groups, with dietary examples. Group Meaning Typical evidence pattern Diet-related examples Group 1 Carcinogenic to humans Sufficient evidence in humans, or strong mechanistic evidence in exposed humans Processed meat; alcoholic beverages; acetaldehyde associated with alcohol consumption; aflatoxins; Chinese-style salted fish Group 2A Probably carcinogenic to humans Usually limited human evidence plus sufficient animal or strong mechanistic evidence Red meat; very hot beverages above 65 °C; acrylamide Group 2B Possibly carcinogenic to humans Limited human evidence, or sufficient animal evidence alone Aloe vera whole-leaf extract; pickled vegetables (traditional Asian) Group 3 Not classifiable Evidence inadequate or does not fit other groups Coffee (reclassified from 2B in 2016, with inverse associations noted for liver and uterine endometrium) Source: compiled from the IARC Monographs Preamble (2019) and IARC Monographs volumes 56, 60, 100E, 100F, 108, 114 and 116; red and processed meat evaluation summarised in Bouvard et al., Lancet Oncology 2015. A former Group 4, "probably not carcinogenic to humans", contained only one agent for many years and was removed in the 2019 revision of the Preamble. Hazard is not risk The single most important thing to understand about IARC classifications is that they describe the strength of evidence that an agent can cause cancer (hazard), not how much cancer it causes at typical exposures (risk). Processed meat and tobacco smoking are both Group 1, because the evidence that each can cause cancer is sufficient. They are not equally dangerous. Tobacco causes a large proportion of lung cancers and many other cancers, with very large relative risks; processed meat raises colorectal cancer risk moderately. When IARC evaluated red and processed meat in October 2015, the working group reported that each 50 g portion of processed meat eaten daily was associated with an increase in colorectal cancer risk of about 18 per cent, and that for red meat, if the association were causal, each 100 g per day would raise risk by about 17 per cent. These figures are relative risks. Their meaning for an individual depends on the baseline risk. Worked example (illustrative). Suppose, purely for arithmetic, that a population has a lifetime colorectal cancer risk of 5 per cent among people who eat no processed meat. A relative risk of 1.18 applied to that baseline gives 5 × 1.18 = 5.9 per cent. The absolute increase is 0.9 percentage points, or roughly one additional case per 110 people over a lifetime at that level of intake. That is meaningful across a population of millions, which is why public health bodies act on it, but it is a very different message from "processed meat is as dangerous as cigarettes". When you explain risk to a patient or write for the public, give absolute as well as relative figures, and state the baseline you have assumed. Two further distinctions help. First, IARC classifications do not account for dose: the hazard identification is the same whether exposure is trivial or heavy. Separate risk assessment, often by national food safety authorities, considers intake levels. Second, IARC does not make dietary recommendations; WCRF/AICR and national bodies do. So a well-constructed essay would say something like: "IARC classifies processed meat as Group 1, indicating sufficient evidence that it causes colorectal cancer in humans; the WCRF/AICR CUP independently judged the evidence convincing and recommends eating little, if any, processed meat." Absolute risk, attributable fractions and public health Public health bodies also estimate the proportion of cancers attributable to modifiable factors. Estimates for high-income countries generally suggest that a substantial minority of cancers, often cited in the region of a third to four in ten, are linked to modifiable risk factors including tobacco, excess body weight, alcohol, diet, physical inactivity, infections and ultraviolet radiation. Tobacco remains the largest single contributor in most such analyses, with excess body weight usually among the next largest. These population attributable fractions depend on both the strength of the association and how common the exposure is. They are estimates built on assumptions, and you should quote them with that caveat and with the source. How to write about evidence in exams and case studies Several habits distinguish strong student work. • Name the grade and the grading body, not just "evidence suggests". • Distinguish the cancer site and, where relevant, the subtype or menopausal status. • Distinguish relative from absolute risk, and hazard from risk. • Say what the study design can and cannot show. "In prospective cohort studies, higher intake was associated with lower risk, but residual confounding by overall lifestyle cannot be excluded" is precise and honest. • Connect the epidemiology to a plausible mechanism from Chapter 1, and say whether that mechanism has been demonstrated in humans or only in models. • Avoid the phrase "cancer-fighting food". It merges prevention with treatment and overstates the evidence for almost every food it is applied to. With these tools in hand, we can examine the specific dietary exposures that damage DNA. Chapter 3: Dietary Carcinogens and DNA Damage The first half of the prevention pathway concerns things in or produced from food that damage DNA. These are the exposures most closely tied to initiation and to the enabling characteristic of genome instability. They are also the exposures that give rise to the most alarming headlines, so precision is especially important. This chapter works through the four groups that students most need to understand in mechanistic detail: N-nitroso compounds and haem iron, heterocyclic amines and polycyclic aromatic hydrocarbons, aflatoxin, and alcohol. It closes with salt-preserved foods and a note on acrylamide, and a summary table. A general model of chemical DNA damage Before looking at individual agents, it helps to have a single sequence in mind, because almost all dietary genotoxins follow some version of it. Exposure. The compound, or a precursor, is ingested or formed in the body. Metabolic activation. Phase I enzymes, chiefly cytochrome P450 isoforms, convert the compound into an electrophilic intermediate, a molecule hungry for electrons. Detoxification or activation by phase II. Conjugation with glutathione, glucuronic acid, sulfate or acetyl groups may neutralise the electrophile, or, for some compounds, create an even more reactive species. Adduct formation. The electrophile bonds covalently to nucleophilic sites on DNA bases, especially the N7 and O6 positions of guanine and the exocyclic amino groups. Repair or mutation. DNA repair systems, including base excision repair, nucleotide excision repair and direct reversal by O6-methylguanine-DNA methyltransferase (MGMT), remove many adducts. Unrepaired adducts can cause mispairing during replication, fixing a mutation. Selection. If the mutation falls in a gene that confers a growth advantage, such as KRAS, APC or TP53, the cell may begin to form a clone. The factors that determine outcome at each step, such as exposure level, enzyme genotype, repair capacity, and the rate of cell division in the target tissue, are why the same diet does not produce the same risk in everyone. They also explain why tissues with high turnover, such as the colonic epithelium, are vulnerable. Processed meat, N-nitroso compounds and haem iron Processed meat is meat transformed through salting, curing, fermentation, smoking or other processes to enhance flavour or improve preservation. The category includes ham, bacon, salami, sausages containing cured meat, hot dogs and corned beef. Red meat means unprocessed mammalian muscle meat: beef, veal, pork, lamb, mutton, horse and goat. These definitions, used by IARC and WCRF/AICR, matter because students often misclassify fresh pork (red) or chicken nuggets (processed poultry, which was not the focus of the IARC evaluation). Two mechanisms dominate the explanation for why processed meat, and to a lesser extent red meat, increase colorectal cancer risk. N-nitroso compounds (NOCs). Nitrite and nitrate are added to cured meats as preservatives; nitrite inhibits Clostridium botulinum and gives cured meat its pink colour. Under acidic conditions, nitrite forms nitrosating agents that react with secondary amines and amides to produce nitrosamines and nitrosamides. Some are formed during processing and cooking (frying bacon at high temperature is a well-known source), but a substantial fraction is formed endogenously in the gut. Many NOCs, after metabolic activation (for nitrosamines, α-hydroxylation by cytochrome P450), generate alkylating agents that methylate or ethylate DNA. The O6-alkylguanine adduct is particularly mutagenic because it pairs with thymine rather than cytosine, causing G to A transitions. Such mutations are found in KRAS in colorectal tumours, which is mechanistically consistent, although not proof of cause. Haem iron. Red meat is rich in haem, the iron-containing porphyrin of myoglobin. Haem contributes to colorectal carcinogenesis in at least three ways. First, it catalyses endogenous nitrosation, forming nitrosyl haem and nitrosothiols; human feeding studies have shown that faecal NOC levels rise with red meat intake in a dose-dependent way, whereas white meat has little effect. Second, haem catalyses the peroxidation of dietary fats, producing reactive aldehydes such as malondialdehyde and 4-hydroxynonenal that can form DNA adducts and are cytotoxic. Third, in animal models haem damages the surface epithelium of the colon, provoking compensatory hyperproliferation of crypt cells. This third mechanism is promotional rather than initiating: repeated injury and repair increases the number of cell divisions and so the opportunity for mutations to become fixed. The combination explains why processed meat, which contains both haem and added nitrite, has a stronger and more consistent association with colorectal cancer than red meat, and why poultry and fish are not associated in the same way. It also helps explain observations from animal studies that calcium, which binds haem, and certain plant compounds can blunt some of these effects, though translation to human dietary advice is limited. A note on vegetables: vegetables such as beetroot, spinach and lettuce contain far more nitrate than cured meats do, yet vegetables are not associated with colorectal cancer in the same way. The difference is thought to lie in the presence in vegetables of vitamin C and polyphenols, which inhibit nitrosation, and in the absence of haem and amines. This is a good example for an essay on why the food matrix matters more than a single chemical. Heterocyclic amines and polycyclic aromatic hydrocarbons These two groups of compounds are formed not by processing but by cooking, particularly at high temperature, and are linked to cooking methods rather than to meat per se. Heterocyclic amines (HCAs), also called heterocyclic aromatic amines, form when muscle meat (beef, pork, poultry or fish) is cooked at high temperatures, generally by frying, grilling or barbecuing, especially until well done or charred. They arise from reactions between creatine or creatinine, free amino acids and sugars, all of which are abundant in muscle. Among the most studied are PhIP (2-amino-1-methyl-6-phenylimidazo[4,5-b]pyridine) and MeIQx. Their activation follows the general model precisely. In the liver, cytochrome P450 1A2 (CYP1A2) N-hydroxylates the amine group. The N-hydroxy metabolite is then O-acetylated by N-acetyltransferase 2 (NAT2), or O-sulfated, producing an unstable ester that breaks down to a highly reactive arylnitrenium ion. This ion binds mainly to the C8 position of guanine. Genetic variation matters: people who are "rapid acetylators" (NAT2 genotype) with high CYP1A2 activity may generate more of the reactive species, and some studies have suggested that the association between well-done meat and colorectal cancer is stronger in such people, though these gene-diet interaction findings have not been consistent. Polycyclic aromatic hydrocarbons (PAHs) form when organic matter burns incompletely. In cooking, they arise when fat and juices drip onto a flame or hot coals and the resulting smoke deposits PAHs back on the food, and in smoking of meats and fish. They are also present in tobacco smoke and polluted air. The best-studied is benzo[a]pyrene. It is activated in a three-step sequence: CYP1A1 or CYP1B1 forms an epoxide; epoxide hydrolase converts this to a dihydrodiol; a second P450 oxidation produces benzo[a]pyrene-7,8-diol-9,10-epoxide (BPDE), the ultimate carcinogen, which binds the N2 position of guanine. BPDE adducts preferentially occur at particular codons of TP53 that are mutational hotspots in smoking-related lung cancer, which is a striking example of a carcinogen leaving a molecular fingerprint. The epidemiology for HCAs and PAHs specifically is less secure than for processed meat as a whole, partly because it is hard to measure cooking methods and doneness accurately over a lifetime. The WCRF/AICR CUP judged that the evidence linking grilled or barbecued (charbroiled) meat and fish to stomach cancer was limited-suggestive. Practical advice that follows from the mechanism is reasonable and low-cost: avoid charring, turn meat frequently, pre-cook in the oven or microwave to reduce time on the grill, trim fat to reduce dripping, and remove blackened portions. Aflatoxin: a model of gene-environment interaction Aflatoxins are toxins produced by moulds, chiefly Aspergillus flavus and Aspergillus parasiticus, that grow on crops such as maize, groundnuts (peanuts), tree nuts and some spices when they are stored in warm, humid conditions. Aflatoxin B1 is the most potent. Exposure is highest in parts of sub-Saharan Africa and South-East Asia where staple crops are stored without adequate drying, and where regulatory control is limited. The mechanism is one of the clearest in diet-related carcinogenesis: Aflatoxin B1 is absorbed and reaches the liver. Hepatic cytochrome P450 enzymes, mainly CYP1A2 and CYP3A4, oxidise it to aflatoxin B1-exo-8,9-epoxide. The epoxide binds the N7 position of guanine, forming an adduct that can be converted into a more stable, persistent lesion. The adduct causes G to T transversions. In hepatocellular carcinomas from high-exposure regions, a large proportion carry a specific mutation at codon 249 of TP53**, changing arginine to serine (AGG to AGT). This signature mutation is rare where exposure is low. Glutathione S-transferases conjugate and detoxify the epoxide; variation in these enzymes affects susceptibility. Aflatoxin interacts strongly with chronic hepatitis B virus (HBV) infection. Cohort studies in China found that the risk of liver cancer in people with both biomarkers of aflatoxin exposure and HBV infection was far higher than with either alone, consistent with a multiplicative interaction. Hepatitis B causes chronic inflammation and hepatocyte turnover, which increases the chance that aflatoxin adducts become fixed mutations; it may also alter enzyme expression. This is an excellent case study for essays on gene-environment and infection-diet interaction, and it has public health lessons: HBV vaccination and better post-harvest storage both reduce liver cancer burden. WCRF/AICR judged the evidence that aflatoxins cause liver cancer to be convincing. Alcohol and acetaldehyde Alcohol is the dietary exposure with the broadest and most firmly established link to cancer. IARC classifies both alcoholic beverages and ethanol in alcoholic beverages as Group 1, together with acetaldehyde associated with the consumption of alcoholic beverages. WCRF/AICR judged alcohol a convincing cause of cancers of the mouth, pharynx and larynx, oesophagus (squamous cell carcinoma), liver, colorectum and breast (postmenopausal), with evidence for premenopausal breast and stomach judged probable. For several sites there is no clear threshold below which no increase in risk is observed. This is why the 2018 recommendation is, simply, that for cancer prevention it is best not to drink alcohol. Mechanisms are multiple: Acetaldehyde. Ethanol is oxidised to acetaldehyde, mainly by alcohol dehydrogenase (ADH) in the liver, and also by CYP2E1 and by bacteria in the mouth and colon. Acetaldehyde is then oxidised to acetate by aldehyde dehydrogenase, principally mitochondrial ALDH2. Acetaldehyde is genotoxic: it forms DNA adducts, notably N2-ethylidene-deoxyguanosine, as well as DNA-protein and DNA interstrand crosslinks, and it impairs DNA repair. Salivary acetaldehyde can reach high local levels because oral bacteria produce it from ethanol and salivary glands have limited ALDH2, which helps explain the strong upper aerodigestive tract risk, especially in combination with smoking. The ALDH2 natural experiment. A common variant, ALDH2\2*, found in a substantial proportion of people of East Asian ancestry, produces an enzyme with very low activity. Carriers experience facial flushing, nausea and tachycardia after drinking because acetaldehyde accumulates. Heterozygous carriers who nonetheless drink regularly have markedly higher risk of oesophageal squamous cell carcinoma than drinkers with normal ALDH2. Because the genotype is inherited, this is strong evidence that acetaldehyde itself is causal. For practice, it means that a patient who reports alcohol flushing is signalling a higher personal risk from drinking. Oxidative stress and CYP2E1. Chronic heavy drinking induces CYP2E1, which generates reactive oxygen species and lipid peroxidation products that damage DNA. CYP2E1 also activates various procarcinogens, including some nitrosamines. Solvent effect. Ethanol increases the permeability of the oral and oesophageal mucosa to other carcinogens, notably those in tobacco smoke. Alcohol and tobacco together multiply rather than simply add risk for head and neck and oesophageal cancers. Hormonal effects. Alcohol raises circulating oestrogen levels, which contributes to its effect on breast cancer. This effect is seen even at relatively low intakes. Folate and one-carbon metabolism. Alcohol impairs folate absorption and metabolism and can deplete folate. Folate is needed for DNA synthesis and methylation; deficiency causes uracil misincorporation and aberrant methylation. Some studies suggest the alcohol-colorectal and alcohol-breast associations are stronger in people with low folate intake. Liver disease. Heavy drinking causes cirrhosis, which is itself the main precursor of hepatocellular carcinoma through cycles of necrosis and regeneration. For essays, alcohol is the best example you have of a dietary factor that acts through several hallmarks at once, with a genetic natural experiment supporting causality. Salt, salt-preserved foods and stomach cancer Stomach cancer rates have fallen dramatically in high-income countries over the twentieth century, and one explanation is the replacement of salting and pickling by refrigeration for food preservation. High salt intake damages the gastric mucosa, causing inflammation and increased cell turnover, and appears to enhance colonisation by Helicobacter pylori, the major cause of non-cardia gastric cancer and itself a Group 1 carcinogen. Salted and preserved foods may also contain NOCs. WCRF/AICR judged that consumption of foods preserved by salting probably increases stomach cancer risk. Chinese-style salted fish, prepared with less salt and often partially decomposed, contains volatile nitrosamines and is a Group 1 carcinogen for nasopharyngeal carcinoma, particularly when consumed in childhood, in southern China and South-East Asia. Acrylamide and other process contaminants Acrylamide forms when starchy foods such as potatoes, bread and cereals are cooked at high temperature, through the Maillard reaction between the amino acid asparagine and reducing sugars. It is metabolised by CYP2E1 to glycidamide, which forms DNA adducts. IARC classifies acrylamide as Group 2A on the basis of animal and mechanistic evidence. Epidemiological studies of dietary acrylamide have not shown consistent associations with human cancer. Regulatory advice, for example the UK Food Standards Agency's "Go for Gold" message, encourages cooking starchy foods to a golden yellow rather than dark brown. This is a good example of precautionary advice based on hazard where human risk at dietary levels remains uncertain; present it as such. Summary of the main dietary genotoxins Table 2 brings the exposures in this chapter together, matching each to its principal mechanism and the cancer sites with the strongest evidence. It is a revision aid, not a substitute for the mechanisms described above. Table 2. Main dietary genotoxic exposures, mechanisms and principal sites. Exposure Key agent(s) Principal mechanism Main site(s) and WCRF/AICR grade Processed meat NOCs, haem, nitrite Alkylating DNA adducts; haem-driven nitrosation, lipid peroxidation, mucosal hyperproliferation Colorectum: convincing increase Red meat Haem iron Endogenous nitrosation; lipid peroxidation; cytotoxicity Colorectum: probable increase High-temperature cooked meat HCAs, PAHs CYP1A2/NAT2 and CYP1A1/1B1 activation to guanine adducts Stomach (grilled meat and fish): limited-suggestive Aflatoxin-contaminated foods Aflatoxin B1 CYP-derived epoxide; TP53 codon 249 mutation; synergy with HBV Liver: convincing increase Alcoholic drinks Ethanol, acetaldehyde DNA adducts, ROS via CYP2E1, solvent effect, raised oestrogen, folate disruption Mouth, pharynx, larynx, oesophagus (squamous), liver, colorectum, postmenopausal breast: convincing Salt-preserved foods Salt, NOCs Mucosal damage, inflammation, enhanced H. pylori effects Stomach: probable increase Source: compiled from the WCRF/AICR Third Expert Report (2018) and Continuous Update Project cancer-site reports; IARC Monographs volumes 56, 96, 100E, 100F and 114. Practical translation Mechanistic knowledge should shape proportionate advice. For prevention, it supports limiting processed meat to little or none, keeping red meat moderate, avoiding charring, minimising alcohol, and choosing foods stored safely. It does not support telling people that a single grilled steak or a glass of wine is dangerous, nor that nitrate-rich vegetables should be avoided. And, as the treatment chapters will show, none of this is a priority for a patient who is losing weight during chemotherapy. In that situation, a ham sandwich that the patient will actually eat is better nutrition than a lentil salad they cannot face. Hashtags: #TheCancerDietConnection #NutritionalOncology #CancerNutrition #CancerPrevention #OncologyNutritionSupport #HallmarksOfCancer #Carcinogenesis #Apoptosis #Angiogenesis #TumourPromotingInflammation #GenomeInstability #ProcessedMeat #RedMeat #NitrosoCompounds #HaemIron #HeterocyclicAmines #PolycyclicAromaticHydrocarbons #Aflatoxin #AlcoholAndCancer #CancerCachexia #MalnutritionScreening #NutritionSupport #TreatmentTolerance #CancerSurvivorship #FutureOfNutritionalOncology
- The Chemistry of Digestion (A Student's Guide to Nutritional Biochemistry)
Download the Book (PDF): Introduction Tom Brody's textbook is the heavyweight champion of biochemical nutrition. Packed with complex molecular structures, receptor kinetics, and intricate metabolic flowcharts, it is designed for deep, graduate-level research. For Level 5 students encountering these complex enzymatic reactions in their second academic year, this book frequently feels like an impenetrable wall of organic chemistry. This book is your strategic shortcut through the chemistry. It is explicitly written to explain Nutritional Biochemistry, breaking down Brody's exhaustive chemical analysis into highly focused, digestible study modules. We isolate the absolute essentials of vitamin and mineral coenzymes, lipid metabolism, and genetic expression. By translating the heavy molecular biology into clear, conceptual frameworks, this guide ensures your biochemistry exams test your understanding, not just your ability to memorize formulas. This guide is an educational study aid for students of biochemistry and nutrition science; it is not clinical advice, and nothing in it should be used to diagnose, treat, or manage any condition in a real person. The book you are studying Nutritional Biochemistry was written by Tom Brody, then in the Department of Nutritional Sciences at the University of California, Berkeley. The second edition was published by Academic Press in San Diego in 1999, running to roughly a thousand pages. It sits at an unusual point in the literature. Most nutrition textbooks are organised around foods, populations, and recommended intakes, with biochemistry appearing as a thin supporting layer. Most biochemistry textbooks are organised around pathways and enzymes, with nutrients appearing only as substrates. Brody's book refuses that division. It treats each nutrient as a chemical problem: what molecule enters the body, how the body converts it into something catalytically or structurally useful, what reactions depend on it, what breaks when it is absent, and what evidence establishes each of those claims. The structure follows that logic. Early chapters deal with the classification of biological structures, digestion and absorption, and the material that resists digestion. The middle chapters handle the regulation of energy metabolism and energy requirement, then lipids, obesity, and protein. Two very large chapters then cover the vitamins and the inorganic nutrients, and a final chapter addresses diet and cancer. Appendices deal with nutrition methodology, molecular techniques, and the language of epidemiology. That last group of appendices is easy to skip and expensive to skip, because Brody's habit throughout is to justify a claim by describing the experiment that produced it. That habit is the source of both the book's authority and its difficulty. Brody rarely tells you that thiamin deficiency causes beriberi and moves on. He tells you which enzymes require thiamin pyrophosphate, what happens to pyruvate and to the branched-chain keto acids when those enzymes lose their cofactor, how erythrocyte transketolase activation is measured, and which human studies established the dose-response relationship. For a researcher this is exactly right. For a second-year student with a deadline, it can read as an undifferentiated wall of detail in which the load-bearing facts are not visually distinguishable from the supporting ones. What this guide does The controlling idea here is simple, and if you take nothing else from this book, take this: nutritional biochemistry becomes learnable the moment you stop memorising structures and start asking three questions of every nutrient. What specific reaction does it make possible? What accumulates or fails when it is absent? How is the whole system regulated when supply changes? A structure you cannot draw is rarely what an examiner is testing. A reaction you cannot name, a deficiency you cannot explain mechanistically, and a regulatory loop you cannot describe are exactly what they are testing. Those three questions do real work. They convert the vitamin chapters from a list of molecules into a short catalogue of chemistry: thiamin handles the removal of carbon dioxide from alpha-keto acids; biotin carries carbon dioxide onto substrates; pyridoxal phosphate handles almost everything that happens to an amino group; folate carries one-carbon units at three oxidation levels; cobalamin performs two reactions in humans and no others. Once you can say that much, the deficiency syndromes are largely deducible rather than memorable, and the biomarkers follow from the blocked reactions. The second thing this guide does is treat the age of the source honestly. Brody's chemistry is durable. Glycolysis has not changed since 1999, the urea cycle has not changed, and the carboxylation chemistry of biotin and vitamin K is exactly as he describes it. But the regulatory and policy layers of the field have moved substantially, and a student who reproduces a 1999 regulatory account in a 2026 examination will lose marks for accuracy, not for effort. So this guide adopts a deliberate then and now approach. Where the textbook's account still stands, we say so and explain it. Where the field has moved, we mark the shift explicitly and give the current position with a verifiable source. The major shifts are worth naming at the outset so you can watch for them. Iron homeostasis was reorganised around the peptide hormone hepcidin, which was not known when the book was written and which now explains absorption control, the anaemia of inflammation, and the hereditary iron-overload disorders in a single mechanism. Vitamin D receptor biology expanded far beyond calcium, and then a decade of large randomised trials largely failed to convert that biology into clinical benefit, which is itself an important lesson about how mechanistic plausibility relates to evidence. Mandatory folic acid fortification began in the United States in 1998, and its outcomes and its residual controversies are now measurable rather than predicted. The gut microbiome went from a footnote about colonic fermentation to a research field in its own right. Nutrient regulation of gene expression, which Brody treats as an emerging topic, is now a mature discipline that includes epigenetic mechanisms he could not have described. Reference intakes have been revised repeatedly, and even the ATP yield of glucose oxidation, printed as a confident integer in older textbooks, is now given as a range. How the material is arranged The guide opens with methodology, because you cannot evaluate a nutritional claim without knowing how the claim was generated, and because Brody's own arguments constantly rest on experimental design. From there it follows food through the body: digestion and absorption first, then the central pathways of energy metabolism, then carbohydrate, lipid, and protein metabolism in turn, each treated as a system with inputs, outputs, and control points rather than as a sequence of arrows to be memorised. The vitamins occupy two chapters, split as the chemistry splits, with the water-soluble vitamins treated as coenzymes attached to named reactions and the fat-soluble vitamins treated as a more heterogeneous group in which two members are effectively hormones, one is a lipid-phase antioxidant, and one is a cofactor for a single unusual post-translational modification. Minerals follow, organised by the chemistry that makes each element useful: redox cycling, Lewis acidity, structural coordination, or electrochemical gradient. The book closes with two integrative chapters, one on how nutrients reach the genome and one on oxidative chemistry and its defences, because these are the areas where a 1999 text is least complete and where current examinations are most likely to probe. Structures and pathways throughout are written as numbered step sequences in prose. This is deliberate. A pathway you can recite as a series of chemical events, each with a named enzyme and a named cofactor, is a pathway you can reconstruct under exam conditions; a pathway you have only ever seen as a picture tends to vanish the moment the picture is taken away. Learn them as sentences, and the diagrams in Brody will become confirmations of what you already know rather than objects to be copied. Chapter 1: How We Know: The Tools of Nutritional Biochemistry Brody places nutrition methodology in an appendix, which is a reasonable editorial decision and a poor pedagogical one. Almost every substantive claim in the main text rests on a method described there: a nutrient is called essential because of what happened in a depletion experiment, a requirement is quoted because of a balance study, a coenzyme function is asserted because of an enzyme assay. If you read the vitamin chapters without knowing what an erythrocyte glutathione reductase activation coefficient is, you will experience the numbers as arbitrary. If you know what it measures, the same numbers become an argument. So this chapter comes first. It is not a detour. It is the grammar of everything that follows. Establishing essentiality and estimating requirement Deprivation and repletion The oldest and most decisive method in nutritional science is deprivation. Remove a single substance from an otherwise complete diet, observe a specific deterioration, restore the substance, and observe recovery. The logic is the logic of the classical genetics experiment, and it is why the discovery period of the vitamins, roughly 1910 to 1950, produced such rapid and durable results. The method has strict conditions. The diet must be complete in every other respect, which is why the development of purified and chemically defined diets was itself a scientific advance rather than a technical convenience. The deficiency must produce a reproducible syndrome rather than general decline. And repletion must reverse it, which distinguishes a nutritional deficiency from a toxic effect of the experimental diet. Brody is careful about this last point, and you should be too: a rat that stops growing on a modified diet has not necessarily been deprived of anything, because the diet may simply be unpalatable or poisonous. Human depletion-repletion studies apply the same logic under ethical constraint. Volunteers consume a diet deliberately low in one nutrient, under controlled conditions, while biochemical and functional markers are tracked. The nutrient is then restored in graded doses until the markers normalise. This design gives something no observational study can give: the intake at which a measurable function actually fails, and the intake at which it is restored. Much of what is known about vitamin C, thiamin, riboflavin, and vitamin B6 requirements comes from studies of exactly this type, many conducted in metabolic wards in the mid-twentieth century under conditions no ethics committee would now approve. That is one reason some requirement estimates have been slow to change: the foundational experiments cannot easily be repeated. Nutrient requirements in humans are therefore built on a mixture of evidence, and modern reference intakes acknowledge this explicitly. An Estimated Average Requirement is the intake meeting the needs of half of a defined group; a Recommended Dietary Allowance is set roughly two standard deviations above it, covering about 97 to 98 per cent; an Adequate Intake is used when the data are too thin to construct an EAR and is essentially an informed observation of what healthy people consume; and a Tolerable Upper Intake Level marks where risk of harm begins to rise. When a nutrient carries an AI rather than an RDA, as pantothenic acid, biotin, and vitamin K do, that label is telling you the evidence base is weak, which is useful information in itself. Balance studies and their limits A balance study measures everything going in and everything coming out. For a mineral, that means dietary intake against faecal, urinary, and where relevant dermal losses over a defined period. Positive balance implies accumulation, negative balance implies depletion, and zero balance implies that intake matches obligatory losses. The requirement is then inferred as the intake at which balance is achieved. Nitrogen balance is the canonical example and the one you should be able to discuss. Dietary nitrogen is measured, usually by multiplying protein intake by 0.16 since protein averages about 16 per cent nitrogen. Losses are summed: urinary nitrogen, dominated by urea, plus faecal nitrogen, plus miscellaneous losses through skin, hair, and nails. An adult in stable weight on adequate intake should be in approximate zero balance. Growth, pregnancy, and recovery from injury produce positive balance. Starvation, insufficient protein, uncontrolled diabetes, sepsis, and immobilisation produce negative balance. The technique looks rigorous and has a systematic flaw that you should be prepared to name. Intake is measured by weighing food, which people and investigators tend to do generously, while losses are collected, which is always incomplete. Both errors push in the same direction, so balance studies systematically overestimate retention. The consequence is real: protein requirement estimates derived from nitrogen balance are widely regarded as somewhat low, and estimates from newer isotopic methods, particularly the indicator amino acid oxidation technique, run higher. For a student, the useful lesson is not the number but the reasoning. When a measurement error is systematic rather than random, averaging over more subjects does not help. Tracers and functional assays Isotope tracers Isotopes let investigators follow a specific atom rather than a bulk quantity, and they turned nutrition from an accounting discipline into a kinetic one. Two classes matter. Stable isotopes, such as carbon-13, nitrogen-15, deuterium, and oxygen-18, are non-radioactive and are detected by mass spectrometry. They can be given to children and pregnant women, which is why they dominate modern human work. Radioactive isotopes such as carbon-14, iron-59, and cobalt-57 are detected by their decay and offer extraordinary sensitivity at very low chemical doses, which is why they remain valuable for nutrients present in trace quantities, though human use is now tightly restricted. Three applications are worth understanding in detail. First, absorption and bioavailability: a labelled dose of iron or zinc is administered, and appearance in blood or whole-body retention is measured, which is how the profound effect of phytate and the enhancing effect of ascorbate on non-haem iron absorption were quantified. Second, pool size and turnover: a labelled tracer is given, allowed to equilibrate with the body's pool of that substance, and the resulting dilution reveals the pool size, while the rate at which label disappears reveals turnover. Third, pathway tracing: a substrate labelled at a specific carbon position is followed into products, which is how the fate of individual carbons in glycolysis and the citric acid cycle was established, and how it became possible to show that the carbons of a labelled fatty acid do not appear in newly synthesised glucose in any net sense. The doubly labelled water method deserves separate mention because it resolved a long-standing problem. A subject drinks water labelled with both deuterium and oxygen-18. Deuterium leaves the body as water only. Oxygen-18 leaves as both water and carbon dioxide, because bicarbonate exchanges oxygen with body water. The difference between the two elimination rates gives carbon dioxide production, and therefore total energy expenditure, in a free-living person over one to two weeks. This method demonstrated conclusively that self-reported dietary intake is substantially under-reported, typically by a fifth or more and considerably more in people with obesity. Any epidemiological finding based on self-reported intake has to be read with that in mind. Enzyme assays and functional biomarkers The most elegant tools in the vitamin sections are the functional assays, and they follow directly from the coenzyme concept. If a vitamin serves as a cofactor for a specific enzyme, then in deficiency that enzyme circulates in its apoenzyme form, lacking its cofactor. Take a sample of the patient's red cells, measure the enzyme's activity as it stands, then add the cofactor in vitro and measure again. In a replete person, adding cofactor changes little because the enzyme is already saturated. In a deficient person, activity jumps. The ratio of stimulated to basal activity, called the activation coefficient, is a direct functional measure of nutritional status rather than a measure of circulating concentration. Three of these are standard and you should know all three. Erythrocyte transketolase activity, stimulated by thiamin pyrophosphate, assesses thiamin status. Erythrocyte glutathione reductase activity, stimulated by flavin adenine dinucleotide, assesses riboflavin status. Erythrocyte aspartate aminotransferase activity, stimulated by pyridoxal phosphate, assesses vitamin B6 status. The principle extends beyond activation coefficients to metabolite accumulation. When a reaction is blocked for want of a cofactor, the substrate accumulates and is often excreted. Elevated plasma homocysteine indicates impaired remethylation and therefore implicates folate, cobalamin, or riboflavin, while elevated methylmalonic acid is more specific to cobalamin because it reflects the failure of a distinct, cobalamin-dependent reaction. The urinary excretion of xanthurenic acid after a tryptophan load reflects impaired pyridoxal phosphate function. These are not laboratory curiosities; they are the reason a serum concentration is a weaker piece of evidence than a functional test, and they explain why vitamin status is now assessed with panels rather than single numbers. Because these methods differ so much in what they can establish, it is worth setting them side by side. Table 1 summarises the question each design answers and where each one characteristically fails. Table 1. What each research design can and cannot establish. Design Question it answers Characteristic weakness Animal depletion Is the nutrient essential, and what fails without it Species differences in requirement and pathway Human depletion-repletion What intake restores function Small samples, ethical limits, short duration Balance study What intake matches obligatory losses Systematic overestimation of retention Isotope tracer Absorption, pool size, turnover, pathway flux Cost, expertise, tracer may not mirror bulk nutrient Prospective cohort Association with disease over years Confounding, measurement error in intake Randomised trial Does changing intake change outcome Dose, duration, and baseline status may be wrong Animal models and the hierarchy of human evidence Animal models permit control that human work cannot approach: defined diets, defined genetics, controlled environment, tissue sampling, and the ability to run a deficiency to its end point. They also permit genetic manipulation, and the knockout mouse has become a standard instrument for establishing what a nutrient-related protein actually does in a whole organism. The limits are specific rather than general, and naming them precisely is more useful than a vague warning about species differences. Most mammals synthesise ascorbate from glucose and therefore have no dietary vitamin C requirement at all; humans, other primates, and guinea pigs cannot, because of a defect in the gene for L-gulonolactone oxidase, which is why the guinea pig became the scorbutic model. Rats practise coprophagy and therefore obtain B vitamins from their own faeces, which can mask a dietary deficiency unless the animals are housed on mesh. Rats have no gallbladder. Rodent lipoprotein metabolism is dominated by high-density lipoprotein rather than low-density lipoprotein, which is why so much atherosclerosis work moved to genetically modified strains. Rodent basal metabolic rates, and therefore vitamin requirements per unit body mass, are far higher than human. None of this makes animal work invalid. It makes each extrapolation a claim that has to be defended on its own terms. Human evidence forms a rough hierarchy, though it should be read as a hierarchy of question-fit rather than a hierarchy of virtue. Case reports and case series generate hypotheses and occasionally settle them outright, as when a nutrient deficiency appears in patients on incomplete parenteral nutrition and resolves on supplementation; the discovery of a human requirement for selenium and for chromium owes something to exactly this. Ecological studies compare populations and are vulnerable to the ecological fallacy. Case-control studies are efficient for rare outcomes and vulnerable to recall bias, which is acute in nutrition because people reconstruct past diets through the lens of present illness. Prospective cohorts measure intake before disease occurs, removing recall bias but leaving confounding intact. Randomised controlled trials remove confounding by design and are the only design that can establish that changing intake changes outcome. Why good trials produce disappointing results A recurring pattern in nutrition, and one that has shaped the field since Brody wrote, is the strong observational association that fails in trial. Beta-carotene intake tracked with lower lung cancer risk across many cohorts; two large trials, the Alpha-Tocopherol Beta-Carotene study in Finnish male smokers and the Beta-Carotene and Retinol Efficacy Trial in smokers and asbestos-exposed workers, found more lung cancer in the supplemented groups, and CARET was stopped early for that reason. Vitamin E showed similar observational promise and similar trial failure. Vitamin D, discussed at length later in this guide, followed the same arc. There are several honest explanations, and a strong answer in an examination distinguishes them rather than reaching for confounding alone. Confounding is real: people with high intakes of a nutrient differ systematically from those without, in smoking, exercise, income, and health-seeking behaviour, and statistical adjustment is imperfect. Reverse causation matters: early disease alters both appetite and circulating nutrient concentrations, so a low level may be a consequence rather than a cause, which is a particularly strong candidate explanation in the vitamin D literature. Dose and form matter: an isolated high-dose synthetic supplement is not a food matrix, and the beta-carotene trials delivered pharmacological doses of a single carotenoid in a way no diet does. Baseline status matters enormously: supplementing a replete population cannot show benefit, so a null trial in well-nourished volunteers does not establish that the nutrient is unimportant, only that adding more of it to people who already have enough does nothing. And duration matters: a five-year trial may be testing a process that operates over decades. Hold all five explanations in view. The lazy reading of the antioxidant trials is that observational nutrition is worthless. The accurate reading is that observational studies and trials answer different questions, and that a nutrient's necessity for normal function is entirely compatible with supplementation being useless in people who are not deficient. Judging a piece of evidence When you meet an unfamiliar nutritional claim, work through a fixed set of questions rather than reacting to the headline. What was the population, and in particular what was its baseline status for the nutrient in question? What was the exposure: a food, a nutrient within a food matrix, or an isolated supplement, and at what dose relative to dietary intakes? What was the comparator, since a trial against placebo and a trial against a lower dose answer different questions? What was the outcome: a disease event, a surrogate marker, or a biochemical index, and how tightly does that surrogate track the outcome anyone actually cares about? How long did the study run relative to the biology it claims to influence? And how large is the effect relative to its confidence interval, which is a more informative question than whether a p-value crossed a threshold. Two further habits are worth acquiring early. First, look for dose-response, because a graded relationship across the exposure range is much harder to explain by confounding than a simple difference between the top and bottom groups. Second, look for coherence across designs: a claim supported by a plausible mechanism, an animal model, a functional human study, consistent cohort data, and a trial is in a different category from one supported by any single line alone. Brody's own argumentative style, which is to assemble exactly this kind of convergent case before committing to a conclusion, is worth imitating in your own writing. Key Takeaways · Essentiality is established by depletion and repletion; requirement is estimated by balance, isotope, and functional studies, each with a known direction of error. · Nitrogen balance systematically overestimates retention because intake is over-recorded and losses are under-collected. · Doubly labelled water showed that self-reported energy intake is substantially under-reported, which weakens all intake-based epidemiology. · Enzyme activation coefficients measure function rather than concentration: transketolase for thiamin, glutathione reductase for riboflavin, aspartate aminotransferase for vitamin B6. · Animal model limits are specific, not general: no vitamin C requirement in most rodents, coprophagy in rats, and rodent lipoproteins dominated by HDL. · Observational and randomised evidence answer different questions; supplementing replete people is the most common reason a mechanistically sound hypothesis fails in trial. Review Questions 1. Explain why a balance study tends to overestimate nutrient retention, and state whether this makes derived requirements too high or too low. 1. Describe the principle of an enzyme activation coefficient and name the specific assay used for thiamin, riboflavin, and vitamin B6 status. 2. Explain how doubly labelled water measures total energy expenditure, identifying the role of each isotope. 3. Give three specific reasons why the beta-carotene trials in smokers contradicted the observational evidence, distinguishing confounding from reverse causation. 4. Why is an Adequate Intake rather than a Recommended Dietary Allowance assigned to biotin, pantothenic acid, and vitamin K, and what does that tell you about the underlying evidence? 5. A guinea pig and a rat are both placed on a vitamin C-free diet. Predict the outcome in each and justify your answer biochemically. Chapter 2: Digestion and Absorption: The Chemistry of the Gut Nothing in nutritional biochemistry happens until a molecule crosses an epithelial cell. The gut is where the chemistry of food meets the chemistry of the body, and it does its work by a strategy that is simple to state and elaborate to execute: hydrolyse polymers to monomers, solubilise what is not water-soluble, and transport the products across a membrane that is deliberately hard to cross. Brody organises this material around the sequence of the tract, and that is the right way to learn it, provided you attach to each compartment the specific chemical problem it solves. The stomach is not simply a mixing vessel; it is an acid denaturation chamber that also initiates protein hydrolysis. The duodenum is not simply a mixing point; it is where an alkaline environment and a detergent system are imposed on an acidic, immiscible mixture so that enzymes can work. The brush border is not simply a surface; it is a catalytic membrane. The tract as a sequence of chemical compartments Digestion begins in the mouth with mechanical disruption and two enzymes. Salivary alpha-amylase, historically called ptyalin, hydrolyses internal alpha-1,4 glycosidic bonds in starch; it is inactivated by gastric acid, so its contribution is modest and confined to the bolus interior. Lingual lipase, secreted by glands at the base of the tongue, has an acidic pH optimum and therefore continues working in the stomach. It is of limited importance in adults and considerable importance in the neonate, whose pancreatic lipase output is immature and whose diet is largely milk fat. The stomach contributes hydrochloric acid, secreted by parietal cells via the gastric hydrogen-potassium ATPase, which brings luminal pH to roughly 1.5 to 3.5. This acid performs four distinct jobs. It denatures dietary proteins, unfolding them so that peptide bonds become accessible. It converts pepsinogen, secreted by chief cells, into active pepsin by autocatalytic cleavage. It reduces ferric iron to the more soluble ferrous form and releases food-bound minerals. And it provides a partial antimicrobial barrier. Pepsin is an endopeptidase with a preference for bonds adjacent to aromatic and large hydrophobic residues, and it produces large peptides rather than amino acids. Gastric lipase supplements lingual lipase. Parietal cells also secrete intrinsic factor, a glycoprotein required for the absorption of vitamin B12 in the ileum, which is why gastric pathology produces a haematological disease. When acidic chyme enters the duodenum it triggers the release of two hormones from enteroendocrine cells. Secretin, released in response to acid, stimulates pancreatic ductal cells to secrete bicarbonate, raising luminal pH toward 6 to 7 so that pancreatic enzymes can function. Cholecystokinin, released in response to fatty acids and amino acids, stimulates pancreatic acinar cells to secrete enzymes and stimulates the gallbladder to contract. Both hormones slow gastric emptying, which matches delivery to processing capacity. Luminal and brush-border enzymology The pancreas supplies the bulk of digestive catalysis, and its enzymes fall into recognisable families. For carbohydrate, pancreatic alpha-amylase hydrolyses internal alpha-1,4 bonds but cannot touch alpha-1,6 branch points or terminal bonds. Its products are therefore not glucose but maltose, maltotriose, and alpha-limit dextrins, which are small branched oligosaccharides. For protein, the pancreas secretes endopeptidases and exopeptidases as inactive zymogens. Activation follows a strict cascade: enteropeptidase, an enzyme anchored to the duodenal brush border, cleaves trypsinogen to trypsin; trypsin then activates more trypsinogen and also chymotrypsinogen, proelastase, and the procarboxypeptidases. This arrangement confines activation to the lumen and protects the pancreas, and its failure is the basis of pancreatitis. Each endopeptidase has a characteristic specificity: trypsin cleaves after the basic residues lysine and arginine, chymotrypsin after aromatic residues, and elastase after small neutral residues such as alanine and glycine. Carboxypeptidase A and B are exopeptidases that remove residues from the carboxyl terminus, A preferring aromatic and aliphatic residues and B preferring basic ones. For lipid, pancreatic lipase hydrolyses triacylglycerol at the sn-1 and sn-3 positions, yielding two free fatty acids and 2-monoacylglycerol. It cannot work at an oil-water interface coated with bile salts unless colipase, secreted as procolipase and activated by trypsin, anchors it there. Phospholipase A2 removes the sn-2 fatty acid from phospholipids, and cholesterol esterase hydrolyses cholesteryl esters. The brush border completes the job. Its enzymes are integral membrane proteins with their catalytic domains facing the lumen, which places the final hydrolysis immediately adjacent to the transporters that will carry the products away. Sucrase-isomaltase is a single polypeptide with two active sites, hydrolysing sucrose and the alpha-1,6 bonds of limit dextrins. Maltase-glucoamylase removes glucose units from the non-reducing ends of maltose and short oligosaccharides. Lactase-phlorizin hydrolase splits lactose into glucose and galactose, and its post-weaning decline in most of the world's adults is the basis of lactose non-persistence, with the persistence phenotype arising from regulatory variants upstream of the lactase gene that arose independently in European and several African populations. Trehalase handles trehalose. A family of brush-border aminopeptidases and dipeptidases completes protein digestion, though a substantial fraction of peptide hydrolysis occurs inside the enterocyte after the peptide has been transported. Surface area, transit, and the limits of absorption The small intestine is built to maximise contact between luminal contents and epithelium. Three levels of folding do this. The circular folds of the mucosa, the plicae circulares, multiply the area of a simple cylinder several-fold. Villi, finger-like projections roughly half a millimetre to a millimetre long, multiply it again by an order of magnitude. Microvilli on the apical surface of each enterocyte, packed at around two hundred thousand per square millimetre and supported by actin filaments, multiply it again. The result is an absorptive surface conventionally estimated at tens of square metres, and more importantly a surface that carries the brush-border enzymes and transporters in the same place, so that the last hydrolytic step and the first transport step occur within nanometres of one another. Each villus is supplied by a capillary network draining to the hepatic portal vein and a central lymphatic lacteal. That dual drainage is the anatomical basis of a rule you will use repeatedly: water-soluble products, including monosaccharides, amino acids, small peptides, minerals, water-soluble vitamins, and short- and medium-chain fatty acids, leave by the portal route and pass through the liver before reaching the general circulation. Long-chain fatty acids re-esterified into chylomicrons leave by the lymphatic route and reach the systemic circulation first, bypassing hepatic first-pass processing. This is why medium-chain triacylglycerols are useful in fat malabsorption: they are hydrolysed more readily, require less micellar solubilisation, and are absorbed portally. Time matters as much as area. Gastric emptying is regulated so that the duodenum is never overwhelmed, and a meal high in fat empties more slowly than one high in carbohydrate, largely through cholecystokinin and the ileal brake, a feedback mechanism in which nutrients reaching the distal small intestine slow proximal transit. Segmentation contractions mix chyme with secretions and bring it repeatedly into contact with the mucosa, while peristaltic waves move it distally. Between meals the migrating motor complex sweeps residue onward, and its disruption contributes to small intestinal bacterial overgrowth. Nutrients that are absorbed slowly, or delivered in quantities exceeding transporter capacity, run out of intestine before they run out of substrate, and what remains passes to the colon. There is also a diffusion barrier that is easy to overlook. Adjacent to the mucosal surface lies an unstirred water layer, a region of relatively static fluid through which any absorbed molecule must diffuse. For water-soluble nutrients this is rarely limiting. For lipids it is decisive, and it is the reason micelles exist: a free fatty acid molecule in bulk lumen has a vanishing probability of crossing that layer, whereas a micelle carrying many such molecules diffuses across it and then dissociates at the acidic microclimate near the membrane, delivering its contents at high local concentration. Bile acids and the solubilisation problem Fat presents a problem that hydrolysis alone cannot solve. Triacylglycerol is insoluble in water, lipase is water-soluble, and the reaction can occur only at an interface. The available interfacial area of a large fat droplet is negligible relative to its volume, so digestion would be impossibly slow without emulsification. Bile acids solve this. They are synthesised in the liver from cholesterol, and the committed and rate-limiting step is the 7-alpha-hydroxylation of cholesterol by cholesterol 7-alpha-hydroxylase, a cytochrome P450 enzyme. The primary bile acids in humans are cholic acid and chenodeoxycholic acid, and they are conjugated in the liver with glycine or taurine, which lowers their pKa so that they remain ionised and therefore water-soluble at intestinal pH. Their molecular architecture is the point: the hydroxyl groups all lie on one face of the rigid steroid nucleus, producing a molecule with a hydrophilic face and a hydrophobic face. Such a molecule sits at an oil-water interface and, above a critical concentration, aggregates into micelles. The sequence in the duodenum runs as follows. First, mechanical agitation plus bile salts disperse the fat mass into an emulsion of small droplets, greatly increasing interfacial area. Second, colipase binds bile salt-coated droplets and provides an anchor for lipase. Third, lipase hydrolyses triacylglycerol to fatty acids and 2-monoacylglycerol. Fourth, these products, together with bile salts, phospholipids, cholesterol, and fat-soluble vitamins, assemble into mixed micelles, which are small enough to diffuse through the unstirred water layer at the mucosal surface. Fifth, at the membrane the micelle releases its cargo, and the lipid products enter the enterocyte, partly by diffusion and partly through membrane proteins including CD36 and members of the FATP family. Sixth, within the enterocyte, fatty acids and monoacylglycerol are re-esterified to triacylglycerol, chiefly by the monoacylglycerol pathway, packaged with apolipoprotein B-48 into chylomicrons, and secreted into lymphatic lacteals, bypassing the hepatic portal vein. Bile acids are then recovered. Roughly 95 per cent are reabsorbed in the terminal ileum by the apical sodium-dependent bile acid transporter and returned to the liver in the portal blood, a circuit called enterohepatic circulation that recycles the pool several times per meal. Loss of the terminal ileum, whether by surgery or Crohn's disease, breaks this circuit and produces both bile acid diarrhoea and fat malabsorption with secondary deficiency of vitamins A, D, E, and K. Bile acid sequestrant drugs exploit the same mechanism deliberately, interrupting reabsorption so that hepatic cholesterol is diverted into new bile acid synthesis. Then and now. Brody describes bile acids essentially as detergents. Since then they have been recognised as signalling molecules. Bile acids are ligands for the nuclear receptor FXR, which mediates feedback inhibition of their own synthesis, and for the membrane receptor TGR5, which influences energy expenditure and the secretion of glucagon-like peptide-1. FXR agonists have entered clinical use in cholestatic liver disease. A detergent that is also a hormone is a good illustration of how far the regulatory layer of this field has moved. Transport across the enterocyte Absorption is not one process. Learn the transporters by the energy source they use. Glucose and galactose cross the apical membrane on SGLT1, a secondary active transporter that moves two sodium ions down their electrochemical gradient for each sugar molecule moved against its own. The sodium gradient is maintained by the basolateral sodium-potassium ATPase, so the energy ultimately comes from ATP. Fructose crosses apically on GLUT5 by facilitated diffusion, which is why fructose absorption is capacity-limited and why large fructose loads reach the colon and ferment. All three sugars leave the cell basolaterally on GLUT2. Amino acids cross on a family of sodium-dependent and sodium-independent carriers with overlapping specificities grouped by side-chain chemistry, which is why a defect in one system produces a characteristic pattern rather than a general failure: in cystinuria the dibasic amino acid transporter is defective, and in Hartnup disorder the neutral amino acid transporter is. Di- and tripeptides are absorbed separately and efficiently by PepT1, a proton-coupled transporter driven by an acidic microclimate at the brush border. This is a point worth remembering, because peptide absorption is quantitatively more important than free amino acid absorption and explains why protein hydrolysates can be absorbed faster than equivalent free amino acid mixtures. Minerals use dedicated systems. Non-haem iron is reduced at the brush border by duodenal cytochrome b and enters on DMT1; haem iron enters by a separate and still incompletely defined route. Calcium crosses transcellularly through TRPV6 under vitamin D control when intake is low, and paracellularly by diffusion when intake is high. Zinc enters on ZIP4, the transporter mutated in acrodermatitis enteropathica. Patterns of malabsorption Malabsorption is a good diagnostic exercise because each pattern points to a specific chemical failure, and working backwards from symptom to mechanism is exactly the skill an examination tests. A defect in a single brush-border disaccharidase produces an osmotic and fermentative syndrome. In lactose non-persistence, undigested lactose remains in the lumen, draws water osmotically, and is fermented in the colon to short-chain fatty acids, hydrogen, carbon dioxide, and methane, producing bloating, flatus, and loose stool. The hydrogen breath test exploits precisely this: hydrogen generated by colonic bacteria is absorbed, carried to the lungs, and exhaled. Congenital sucrase-isomaltase deficiency produces the same pattern with sucrose and starch rather than lactose. A defect in pancreatic secretion produces failure across all three macronutrient classes, but fat malabsorption dominates the clinical picture because there is no salvage pathway for triacylglycerol comparable to colonic fermentation of carbohydrate. Steatorrhoea, weight loss, and deficiency of the fat-soluble vitamins follow. Cystic fibrosis and chronic pancreatitis are the usual causes, and enzyme replacement is the usual treatment. A defect in the mucosa itself, as in coeliac disease, removes surface area indiscriminately. Villous atrophy shortens or obliterates villi, taking brush-border enzymes and transporters with them. Because the damage is typically most severe proximally, nutrients absorbed mainly in the duodenum and proximal jejunum suffer first, which is why iron and folate deficiency are common early findings, while vitamin B12 deficiency, which depends on the distal ileum, is a later development. A defect in bile delivery, whether from biliary obstruction, ileal loss of the bile acid pool, or bacterial deconjugation of bile salts in overgrowth, impairs micelle formation specifically. Carbohydrate and protein digestion proceed normally; fat and fat-soluble vitamin absorption fail. The characteristic finding is steatorrhoea with preserved pancreatic enzyme output. Finally, a defect in a single transporter produces a highly selective lesion, and these inherited disorders are among the most informative experiments in human nutrition. Acrodermatitis enteropathica, caused by mutations in the zinc transporter ZIP4, produces a dermatitis, diarrhoea, and alopecia syndrome that resolves completely on zinc supplementation. Hartnup disorder affects neutral amino acid transport, and because tryptophan is the precursor for endogenous niacin synthesis, affected individuals can present with a pellagra-like rash. In each case a single missing protein isolates one nutrient, and the resulting phenotype tells you what that nutrient does. What escapes digestion, and who eats it Human enzymes cannot hydrolyse beta-1,4 bonds between glucose units, which is why cellulose passes through, nor most of the other linkages in plant cell walls, resistant starch, and many oligosaccharides. Brody devotes a chapter to this material, and his treatment of colonic fermentation is chemically sound: anaerobic bacteria ferment these substrates to the short-chain fatty acids acetate, propionate, and butyrate, plus hydrogen, carbon dioxide, and methane. Butyrate is the preferred fuel of the colonocyte; propionate travels to the liver and is a gluconeogenic substrate; acetate enters the peripheral circulation. Then and now. What has changed is scale and specificity. Culture-independent sequencing, which became practical after Brody wrote, revealed a colonic community of hundreds of species whose collective gene content greatly exceeds the human genome. That community is now implicated in immune development, drug and xenobiotic metabolism, bile acid transformation, the synthesis of vitamin K2 and several B vitamins, and the production of trimethylamine from dietary choline and carnitine, which the liver oxidises to trimethylamine N-oxide. Two cautions belong with this. First, the great majority of microbiome findings are associative, and causal demonstrations in humans remain scarce. Second, the quantitative nutritional contribution of colonic vitamin synthesis in humans is uncertain, because much of it occurs distal to the main sites of absorption. Treat the microbiome as an active research area with a secure core of fermentation chemistry and a large periphery of unsettled claims. Key Takeaways · Each gut compartment solves a defined chemical problem: acid denaturation in the stomach, neutralisation and emulsification in the duodenum, final hydrolysis at the brush border. · Pancreatic proteases are secreted as zymogens and activated by a cascade that begins with brush-border enteropeptidase acting on trypsinogen. · Pancreatic lipase yields fatty acids and 2-monoacylglycerol, requires colipase to work at a bile salt-coated interface, and depends on micelles to cross the unstirred water layer. · Bile acids are amphipathic because their hydroxyl groups lie on one face of the steroid nucleus; about 95 per cent are recovered in the terminal ileum. · Transporters differ in energy source: SGLT1 is sodium-coupled, GLUT5 is facilitative, and PepT1 is proton-coupled, which explains the clinical patterns of specific absorption defects. · The fermentation chemistry Brody describes is still correct; the modern addition is the scale and signalling role of the colonic microbial community, much of which remains associative. Review Questions 1. Set out the zymogen activation cascade in the duodenum and explain why the first step must occur on the brush border rather than in the pancreas. 1. Explain, in numbered steps, how a triacylglycerol molecule in a meal reaches the lymphatic circulation as part of a chylomicron. 2. Why does resection of the terminal ileum cause deficiency of vitamins A, D, E, and K as well as of vitamin B12? 3. Contrast SGLT1, GLUT5, and PepT1 in terms of driving force, and predict the consequence of a large fructose load for each. 4. Name the four functions of gastric hydrochloric acid and predict the biochemical consequences of long-term acid suppression. 5. State what is securely established and what remains uncertain about the nutritional contribution of the colonic microbiota. Chapter 3: Energy Metabolism: Glycolysis to the Proton Gradient Energy metabolism is where most students first lose their footing, usually because they try to memorise three pathways as three lists. The lists are long, the intermediates are unfamiliar, and nothing connects them. The alternative is to learn the system as a single logical sequence with four stages and a currency conversion at the end. The sequence is this. Fuel molecules are broken down to two-carbon acetyl units. Those units are oxidised completely to carbon dioxide, and the electrons removed are handed to two carriers. Those carriers deliver electrons to a membrane-bound chain of complexes that use the energy released to pump protons across a membrane. The resulting electrochemical gradient drives a rotary motor that makes ATP. Everything else is detail hung on that frame. Glycolysis: ten steps and three control points Glycolysis converts one molecule of glucose into two of pyruvate in the cytosol, with a net yield of two ATP and two NADH. It does not require oxygen, which is why it is the only ATP-generating pathway available to the mature erythrocyte, which has no mitochondria, and the dominant one in muscle during maximal effort. Divide it into an investment phase and a payoff phase. In the investment phase, ATP is consumed. Step one: hexokinase, or glucokinase in liver and pancreatic beta cells, phosphorylates glucose to glucose 6-phosphate, using one ATP. This traps the sugar inside the cell, since the phosphorylated form cannot cross the membrane on a GLUT transporter. Step two: phosphoglucose isomerase converts the aldose glucose 6-phosphate to the ketose fructose 6-phosphate. Step three: phosphofructokinase-1 phosphorylates this to fructose 1,6-bisphosphate, using a second ATP. This is the committed step and the principal control point. Step four: aldolase cleaves the six-carbon bisphosphate into two three-carbon units, dihydroxyacetone phosphate and glyceraldehyde 3-phosphate. Step five: triose phosphate isomerase interconverts the two, so that everything proceeds as glyceraldehyde 3-phosphate. In the payoff phase, everything happens twice per glucose. Step six: glyceraldehyde 3-phosphate dehydrogenase oxidises the aldehyde and simultaneously attaches inorganic phosphate, producing 1,3-bisphosphoglycerate and reducing NAD+ to NADH. This is the only oxidation in glycolysis, and the high-energy acyl phosphate it creates is what makes the subsequent ATP synthesis possible. Step seven: phosphoglycerate kinase transfers that phosphate to ADP, producing ATP by substrate-level phosphorylation. Step eight: phosphoglycerate mutase moves the remaining phosphate from carbon three to carbon two. Step nine: enolase removes water, creating phosphoenolpyruvate, a compound with an exceptionally high phosphoryl transfer potential. Step ten: pyruvate kinase transfers that phosphate to ADP, producing a second ATP and pyruvate. Three enzymes are irreversible under cellular conditions and are therefore the control points: hexokinase or glucokinase, phosphofructokinase-1, and pyruvate kinase. These same three steps are bypassed by different enzymes in gluconeogenesis, which is the general rule for opposing pathways and worth stating as such: irreversible steps are where regulation lives, and where the reverse pathway must take a detour. Phosphofructokinase-1 deserves particular attention because its regulation is the clearest example of integrated control in metabolism. It is inhibited by ATP and by citrate, both signals that energy is abundant, and activated by AMP, a signal that it is not. It is inhibited by falling pH, which protects muscle during extreme lactate accumulation. And it is potently activated by fructose 2,6-bisphosphate, a molecule whose sole function is regulatory. That molecule is made and destroyed by a single bifunctional enzyme carrying both a kinase and a phosphatase domain, and which activity dominates depends on the phosphorylation state of the protein. Glucagon, acting through cyclic AMP and protein kinase A, phosphorylates it, activating the phosphatase, lowering fructose 2,6-bisphosphate, and therefore switching the liver away from glycolysis and toward gluconeogenesis. Insulin does the reverse. One small molecule thus converts a hormonal signal into a flux decision. The NADH produced in step six must be reoxidised or glycolysis stops. Under aerobic conditions its electrons enter mitochondria by one of two shuttles. The malate-aspartate shuttle, used in liver, kidney, and heart, regenerates NADH inside the matrix. The glycerol 3-phosphate shuttle, used in skeletal muscle and brain, delivers electrons to FAD instead, which yields less ATP. Under anaerobic conditions, lactate dehydrogenase reduces pyruvate to lactate, regenerating NAD+ and allowing glycolysis to continue. Lactate is not a waste product: it is exported, taken up by other tissues, and reoxidised to pyruvate, or returned to the liver for gluconeogenesis in the Cori cycle. The link reaction and the citric acid cycle Pyruvate enters the mitochondrial matrix on a dedicated carrier and meets the pyruvate dehydrogenase complex, one of the most instructive enzymes in nutrition because it requires five different coenzymes, four of which derive from vitamins. Thiamin pyrophosphate performs the decarboxylation. Lipoic acid accepts and transfers the resulting acetyl group. Coenzyme A, from pantothenic acid, carries the acetyl group away. FAD, from riboflavin, reoxidises the lipoamide. NAD+, from niacin, takes the electrons. The reaction converts pyruvate to acetyl-CoA, releases one carbon dioxide, and reduces one NAD+. It is irreversible, which is the chemical reason that fatty acids, whose catabolism yields acetyl-CoA, cannot be converted to glucose in net terms. The complex is regulated by covalent modification. A dedicated kinase phosphorylates and inactivates it, and is itself stimulated by high ratios of ATP to ADP, NADH to NAD+, and acetyl-CoA to CoA. A phosphatase reactivates it, and is stimulated by calcium, which links contraction in muscle to increased fuel oxidation. Deficiency of thiamin impairs this complex directly, and the accumulation of pyruvate and lactate that follows is the biochemical basis of the neurological and cardiac features of beriberi. The citric acid cycle then oxidises the acetyl group completely. Step one: citrate synthase condenses acetyl-CoA with the four-carbon oxaloacetate to give the six-carbon citrate. Step two: aconitase isomerises citrate to isocitrate. Step three: isocitrate dehydrogenase oxidises and decarboxylates it to the five-carbon alpha-ketoglutarate, producing NADH and one carbon dioxide; this is the principal rate-limiting step, inhibited by ATP and NADH and activated by ADP and calcium. Step four: the alpha-ketoglutarate dehydrogenase complex, structurally and mechanistically analogous to pyruvate dehydrogenase and requiring the same five coenzymes, produces the four-carbon succinyl-CoA, a second NADH, and a second carbon dioxide. Step five: succinyl-CoA synthetase produces GTP or ATP by substrate-level phosphorylation and yields succinate. Step six: succinate dehydrogenase, uniquely embedded in the inner mitochondrial membrane and identical with Complex II of the respiratory chain, oxidises succinate to fumarate, reducing FAD. Step seven: fumarase hydrates fumarate to malate. Step eight: malate dehydrogenase oxidises malate to oxaloacetate, producing a third NADH and regenerating the starting compound. Per acetyl-CoA, the cycle yields three NADH, one FADH2, one GTP or ATP, and two carbon dioxide molecules. Note that the two carbons released as carbon dioxide are not the same two that entered on this turn; they derive from oxaloacetate. Note also that the cycle is amphibolic, not merely catabolic. Citrate is exported for fatty acid synthesis. Alpha-ketoglutarate and oxaloacetate are transaminated to glutamate and aspartate. Succinyl-CoA is used in haem synthesis. When intermediates are withdrawn, they must be replaced by anaplerotic reactions, of which the most important is pyruvate carboxylase, a biotin-dependent enzyme that converts pyruvate to oxaloacetate and is activated by acetyl-CoA. Chemiosmosis and the revised ATP arithmetic The electron transport chain is four complexes and two mobile carriers in the inner mitochondrial membrane. Complex I, NADH dehydrogenase, accepts electrons from NADH and passes them to ubiquinone, pumping four protons out of the matrix. Complex II, succinate dehydrogenase, feeds electrons from FADH2 into ubiquinone without pumping any protons, which is the entire reason FADH2 yields less ATP than NADH. Ubiquinone carries electrons to Complex III, cytochrome bc1, which passes them to cytochrome c and pumps four protons through the Q cycle. Cytochrome c delivers them to Complex IV, cytochrome c oxidase, which reduces molecular oxygen to water and pumps two protons. Complex IV is the enzyme inhibited by cyanide and carbon monoxide, and its copper centres are one reason copper is a required nutrient. Peter Mitchell's chemiosmotic hypothesis, proposed in 1961 and recognised with the Nobel Prize in Chemistry in 1978, explains what happens next. The pumped protons create an electrochemical gradient across the inner membrane, composed of a pH difference and, more importantly in mitochondria, a membrane potential. This proton-motive force is the intermediate between oxidation and phosphorylation; there is no high-energy chemical intermediate, which is why the search for one failed for decades. Protons return through ATP synthase, Complex V, whose membrane-embedded c-ring rotates as they pass. That rotation turns the central gamma subunit inside the catalytic head, and the resulting conformational changes in three catalytic sites drive ADP and phosphate together and release ATP. Paul Boyer and John Walker shared the 1997 Nobel Prize in Chemistry for establishing this binding-change mechanism and the enzyme's structure. Then and now. Older textbooks state that oxidation of one glucose yields 36 or 38 ATP, using whole-number stoichiometries of three ATP per NADH and two per FADH2. Those integers were never measured; they were inferred. Two lines of evidence overturned them. First, the number of protons required per ATP is set by the number of c subunits in the ATP synthase rotor, which in humans is eight, so one full revolution passes eight protons and makes three ATP, giving roughly 2.7 protons per ATP; adding the proton consumed by the phosphate carrier and the adenine nucleotide translocase brings the cost to something closer to four protons per cytosolic ATP. Second, proton pumping was counted per complex: ten per NADH and six per FADH2. Dividing gives approximately 2.5 ATP per NADH and 1.5 per FADH2. These are not exact integers and were never expected to be, because the proton gradient also leaks and drives other transport processes. Modern textbooks therefore quote a range of roughly 30 to 32 ATP per glucose. Table 2 sets out the arithmetic; be prepared to reproduce it and to explain why the older figure was wrong. Table 2. Consensus ATP yield from complete oxidation of one glucose. Stage Products ATP equivalents Notes Glycolysis 2 ATP, 2 NADH (cytosolic) 2 plus 3 to 5 Shuttle-dependent: 1.5 or 2.5 per NADH Pyruvate dehydrogenase 2 NADH 5 Two turns per glucose Citric acid cycle 6 NADH, 2 FADH2, 2 GTP 15 plus 3 plus 2 Per two acetyl-CoA Total 30 to 32 Using 2.5 per NADH, 1.5 per FADH2 Source: P/O ratios of 2.5 for NADH and 1.5 for FADH2, as adopted in standard biochemistry texts following proton-stoichiometry measurements. Two further points complete the picture. Uncoupling proteins, of which UCP1 in brown adipose tissue is the best characterised, allow protons to re-enter the matrix without passing through ATP synthase, so the energy of the gradient appears as heat. This is the mechanism of non-shivering thermogenesis, and it explains why chemical uncouplers such as 2,4-dinitrophenol cause hyperthermia rather than fatigue. And respiratory control couples the whole system to demand: when ATP is consumed, ADP rises, ATP synthase runs faster, protons return more quickly, the gradient falls, and electron transport accelerates to restore it. Oxygen consumption is therefore governed by ATP use, not the other way round. From pathways to requirement Having traced where the ATP comes from, the next question is how much of it a person needs, and this is the point at which biochemistry becomes a measurable whole-body quantity. Brody devotes a full chapter to energy requirement, and the material rewards attention because it is where thermodynamics, methodology, and public health meet. Start with the energy content of food. Bomb calorimetry burns a sample in oxygen and measures the heat released, giving gross energy. The body does not achieve this figure, for two reasons: not all of what is eaten is absorbed, and protein is not oxidised completely, because nitrogen is excreted as urea rather than as oxides of nitrogen. Correcting for both gives metabolisable energy, and the resulting Atwater factors of approximately 4 kilocalories per gram for carbohydrate, 9 for fat, and 4 for protein are the conventional values. Alcohol contributes about 7. These are averages across mixed diets, and the true figure for a specific food can differ appreciably, which is one reason food labels are less precise than their decimal places suggest. Energy expenditure divides into three components. Basal metabolic rate, measured in a thermoneutral environment after an overnight fast and complete rest, accounts for roughly 60 to 70 per cent of total expenditure in a sedentary adult and is largely a function of fat-free mass, which is why men, larger people, and younger people have higher absolute rates. The thermic effect of food, the energy cost of digesting, absorbing, and processing a meal, accounts for around 10 per cent of intake and is proportionally highest for protein, because amino acid deamination, urea synthesis, and peptide bond formation all cost ATP. Physical activity is the most variable component, ranging from a small fraction in a sedentary person to more than the basal rate in a manual worker or endurance athlete. Two measurement approaches exist. Direct calorimetry measures heat output from a person in an insulated chamber; it is accurate, expensive, and rare. Indirect calorimetry infers energy expenditure from oxygen consumption and carbon dioxide production, and it does something direct calorimetry cannot: the ratio of carbon dioxide produced to oxygen consumed, the respiratory quotient, indicates which fuel is being oxidised. Complete oxidation of carbohydrate gives a respiratory quotient of 1.0, because the stoichiometry of glucose oxidation consumes and produces six molecules of each. Fat oxidation gives approximately 0.7, because fatty acids are more reduced and therefore require proportionally more oxygen. Protein gives about 0.8. A mixed diet typically yields 0.85. A value above 1.0 indicates net lipogenesis, since converting carbohydrate to fat releases carbon dioxide without consuming oxygen. For free-living measurement over days to weeks, doubly labelled water remains the reference method. Then and now. Brody's chapter reflects the reference intakes of the 1990s. The framework has since been rebuilt twice. The Institute of Medicine's 2002 and 2005 reports introduced the Estimated Energy Requirement, an intake predicted to maintain energy balance in a person of defined age, sex, weight, height, and physical activity level, deliberately set as an average rather than as an allowance with a safety margin, because unlike a vitamin, excess energy is itself harmful. In 2023 the National Academies of Sciences, Engineering, and Medicine published a further revision, Dietary Reference Intakes for Energy, which updated the prediction equations using pooled doubly labelled water data, including the International Atomic Energy Agency database, and changed how physical activity and body composition are handled. The practical lesson for a student is that any energy requirement figure carries a date, and quoting one without knowing which framework generated it is a common and avoidable error. Key Takeaways · Glycolysis has three irreversible steps, and those steps are both the regulatory points and the reactions gluconeogenesis must bypass. · Fructose 2,6-bisphosphate is a purely regulatory molecule that converts hormonal signals into a decision between glycolysis and gluconeogenesis in liver. · Pyruvate dehydrogenase requires thiamin pyrophosphate, lipoate, CoA, FAD, and NAD+, and is irreversible, which is why fatty acid carbons cannot yield net glucose. · The citric acid cycle is amphibolic; withdrawal of intermediates requires anaplerotic replacement, principally by biotin-dependent pyruvate carboxylase. · Chemiosmosis means there is no high-energy chemical intermediate between oxidation and phosphorylation; the proton-motive force is the intermediate. · The modern yield is about 30 to 32 ATP per glucose, from measured proton stoichiometry, not the older inferred 36 to 38. Review Questions 1. List the three irreversible steps of glycolysis with their enzymes, and explain why each is a regulatory point. 1. Explain how glucagon lowers hepatic glycolytic flux through fructose 2,6-bisphosphate, naming the enzyme and the covalent modification involved. 2. Name the five coenzymes of the pyruvate dehydrogenase complex and their parent vitamins, and predict which reactions fail in thiamin deficiency. 3. Why does FADH2 yield less ATP than NADH? Answer in terms of proton pumping, not in terms of ATP counts. 4. Derive the figure of approximately 2.5 ATP per NADH from proton stoichiometry, and explain why the older figure of three was accepted for so long. 5. Explain how 2,4-dinitrophenol increases oxygen consumption while decreasing ATP synthesis. Hashtags: #TheChemistryOfDigestion #NutritionalBiochemistry #DigestionAndAbsorption #GastricAcid #PancreaticEnzymes #ZymogenActivation #BrushBorderEnzymes #BileAcids #MicelleFormation #Chylomicrons #EnterohepaticCirculation #SGLT1 #GLUT5 #PepT1 #NutrientTransporters #LactoseMalabsorption #FatMalabsorption #GutMicrobiome #ShortChainFattyAcids #Glycolysis #PyruvateDehydrogenase #CitricAcidCycle #ElectronTransportChain #Chemiosmosis #FutureOfNutritionalBiochemistry
- The Convergence Matrix (A Study Guide to Siebel's Digital Transformation)
Download the Book (PDF): Introduction Digital transformation is not driven by a single technology, but by the simultaneous convergence of four distinct vectors: cloud computing, big data, artificial intelligence, and the Internet of Things (IoT). Silicon Valley pioneer Thomas Siebel outlines this technological tsunami with sweeping authority. However, for a student writing a graded strategy paper, separating high-level corporate marketing from rigorous technological and economic frameworks is a major challenge. This study companion delivers the academic rigor you need. It is written to explain Digital Transformation by Thomas M. Siebel, isolating the mechanics of technological convergence and translating them into structured strategic models. We break down the architecture of modern AI-driven platforms and provide clear rubrics for evaluating corporate vulnerability. This guide is built to help your strategic analyses of enterprise disruption meet the highest academic standards of evaluation. The author and the book Few people have watched as many waves of enterprise computing break as Thomas Siebel. He worked at Oracle during its rise in the 1980s, then founded Siebel Systems, which became the dominant vendor of customer relationship management (CRM) software in the late 1990s. Oracle acquired Siebel Systems in a deal announced in 2005 and completed in early 2006. In 2009 he founded a new company, originally aimed at the energy sector as C3 Energy, later renamed C3 IoT and then C3.ai. That company sells software for building and running artificial intelligence applications inside large organizations, and it went public on the New York Stock Exchange in December 2020 under the ticker symbol AI. Digital Transformation: Survive and Thrive in an Era of Mass Extinction was published by RosettaBooks in 2019. It is not Siebel's first book; he had earlier written about selling and electronic commerce during the internet boom, including Cyber Rules and Taking Care of eBusiness. But this is the book in which he lays out a complete theory of the present moment. Its argument, in compressed form, runs as follows. Business is entering a period of upheaval comparable to the great extinction events in the history of life. The cause is the arrival, at the same time, of four mature technologies: elastic cloud computing, big data, artificial intelligence, and the Internet of Things. Together they make it possible to instrument an entire enterprise, store and integrate everything it knows, and apply machine learning to predict and optimize its operations. Firms that rebuild themselves around these capabilities will prosper; firms that do not will disappear. And because the change touches every function at once, the transformation cannot be delegated to the information technology department. It has to be led by the chief executive. The book matters for three reasons. First, it offers a clear, memorable model of why digital change in the late 2010s felt different from earlier technology cycles, and that model is easy to apply in a strategy paper. Second, it contains one of the more detailed accounts written for a general business audience of what an enterprise AI system actually has to do: ingest data from hundreds of sources, reconcile it, run models against it, and push the results into daily operations. Third, it is written by a practitioner with an unusual vantage point, someone who has sold enterprise software to large corporations and government agencies for four decades. That third point cuts both ways. Siebel is not a disinterested observer. When the book appeared he was chairman and chief executive of C3.ai, and many of its central recommendations, above all its advice that companies should adopt a proven, model-driven AI platform rather than assemble their own from open-source components, describe the product his company sells. Several of its case examples are C3.ai customers. None of this makes the book wrong. It does mean that a careful reader has to sort its claims into categories: those that rest on well-established technical and economic reasoning, those that are plausible but unproven, and those that are best read as positioning. The controlling idea of this guide This guide is organized around one proposition. Siebel's most durable contribution is the idea of convergence: the claim that the value of the four technologies lies in their combination, and that the combination changes the economics of an enterprise in ways no single technology could. Everything else in the book, including the extinction metaphor, the architecture, the applications and the leadership agenda, can be read as a consequence of that idea. A student who understands the convergence mechanism can use it to analyze almost any firm, and can also test Siebel's more commercial claims against it rather than accepting them on his authority. We call the analytical tool that follows from this the convergence matrix. It is not Siebel's term. It is this guide's name for a simple discipline: for any firm or industry, ask how far each of the four technologies has penetrated its operations, where they have been combined, and where the combination could produce a change in cost, speed, or knowledge large enough to reorder competition. The matrix appears in developed form in the final chapter, alongside rubrics for assessing a firm's vulnerability to disruption. Those rubrics are the guide's synthesis, drawn from Siebel's argument and from the wider strategy literature. They are not taken from the book, and they should be cited as a study framework rather than as Siebel's own. How the guide is organized The guide moves from the big claim to the machinery and then back to strategy. It begins with Siebel's evolutionary metaphor, what it captures about corporate mortality, and where the analogy strains. It then follows his account of earlier waves of information technology, because his argument depends on showing that the present wave differs in kind and not only in degree. The four technologies each receive their own treatment. Cloud computing is examined as an economic change as much as a technical one, and the discussion brings the market picture up to date: by the third quarter of 2025, three providers accounted for close to two-thirds of enterprise spending on cloud infrastructure. Big data is treated through the problem Siebel considers hardest, which is not storing data but integrating it into a single, usable representation of the enterprise. Artificial intelligence is explained from first principles, and the guide adds what Siebel could not have covered in 2019: the arrival of generative AI and large language models after the public release of ChatGPT in November 2022, and what that shift does and does not change in his argument. The Internet of Things chapter is where convergence becomes concrete, through the widely reported case of the Italian utility Enel and its smart-meter analytics. The middle of the guide takes apart Siebel's model of the AI-driven enterprise: the technology stack he describes, the idea of a model-driven platform, and the build-versus-buy question that sits at the center of his commercial position. It then turns to applications, including predictive maintenance, fraud detection, supply chain optimization and government uses, with the United States Air Force's predictive maintenance program as a documented case. A separate chapter treats the concerns Siebel raises about cybersecurity, privacy, and public policy, and sets them against the regulatory landscape as it stands in 2026. The final chapter turns to Siebel's call for chief executives to lead transformation personally, reconstructs a practical roadmap from his recommendations, and presents the guide's rubrics for judging a firm's exposure to disruption and for weighing the claims of technology vendors, including Siebel's own company. That company's recent history, which includes a period of falling revenue, a change of chief executive in 2025 and Siebel's return to the role in 2026, offers an instructive test of the book's thesis applied to its author's own firm. Each chapter closes with key takeaways and review questions. The takeaways state what should be retained; the questions are designed to be answered in a paragraph or two and to rehearse the kind of reasoning a strong strategy paper requires. A short glossary, a list of further reading and a set of notes on sources complete the guide. Chapter 1: Punctuated Equilibrium and the Mass Extinction Thesis Siebel opens his book with biology rather than technology. His framing device is the idea of punctuated equilibrium, drawn from evolutionary theory, and the related image of mass extinction. The choice is deliberate. A metaphor of gradual improvement would suggest that incumbents have time to adapt. A metaphor of extinction suggests that the rules of survival themselves are about to change, and that the traits which made a species successful in the old environment may be useless, or even fatal, in the new one. Understanding exactly what this metaphor claims, and what it does not, is the first step toward using Siebel's book well. The biological idea and Siebel's use of it In evolutionary biology, punctuated equilibrium describes a pattern in which species remain relatively stable for long stretches and then change rapidly over short intervals. The fossil record also contains a small number of mass extinction events, periods in which a large share of all living species disappeared within a geologically brief window. The best known is the extinction at the end of the Cretaceous period, about 66 million years ago, widely attributed to an asteroid impact, which ended the age of the non-avian dinosaurs and opened ecological space for mammals. The key feature of such events is that survival depends less on how well adapted a species was to the previous environment than on whether it happened to possess traits suited to the new one. Siebel transposes this pattern to the corporate world. For long periods, he suggests, industries reach an equilibrium: a stable set of leading firms, established business models, and familiar competitive rules. Then a shock arrives that resets the environment. He argues that the combination of cloud computing, big data, artificial intelligence and the Internet of Things is such a shock, and that it is producing a wave of corporate disappearance. To support the point he draws on the rapid turnover of the largest companies. A widely circulated statistic, which Siebel uses, holds that roughly half of the companies on the Fortune 500 list in 2000 had dropped off it within the following two decades, through bankruptcy, acquisition or decline. He also points to research on the shrinking average tenure of companies in major stock indexes, a theme developed by the consulting firm Innosight in its studies of the S&P 500. The rhetorical force of the metaphor lies in three implications. The first is speed: extinction events are abrupt relative to the long periods before them. The second is indiscriminateness: size and past dominance do not guarantee survival, and may even be liabilities if they bind a firm to an obsolete model. The third is opportunity: every mass extinction clears space for new forms to flourish, so the same event that destroys incumbents creates the conditions for new leaders. Siebel's subtitle, with its pairing of surviving and thriving, captures this double edge. What the metaphor gets right The extinction framing captures real features of the economics of digital technology, and a strategy paper can defend them on grounds independent of the metaphor. First, digital technologies change the cost structure of entire activities rather than improving them at the margin. When computing power can be rented by the second, when storage costs fall toward zero, and when a statistical model can make millions of predictions a day, activities that used to require armies of analysts or large capital investments become cheap. A change in cost of that magnitude does not simply make existing firms more efficient; it allows entirely different business designs. The classic cases are familiar. Streaming video did not make video rental stores more efficient; it removed the need for them. Blockbuster, which filed for bankruptcy in 2010, is the standard example, and its story is usually told as a failure to recognize that the environment had changed. Second, digital businesses often exhibit increasing returns. A firm with more data can train better models, which attract more users, who generate more data. Network effects and data feedback loops mean that early leads compound. In such markets the gap between winners and losers widens quickly, which is consistent with the abrupt pattern the metaphor predicts. Third, the metaphor correctly identifies that incumbency can be a disadvantage. Clayton Christensen's work on disruptive innovation, set out in The Innovator's Dilemma in 1997, showed how well-managed firms can fail precisely because they listen carefully to their best customers and invest rationally in their existing business. Siebel's point is related but distinct. He is less concerned with low-end entrants creeping upmarket, which is the heart of Christensen's model, and more concerned with a general technology shift that raises the performance bar for every firm in every industry at once. Legacy systems, organizational silos and entrenched processes are, in his account, the traits that made firms fit for the old environment and that now slow their adaptation. Fourth, digital capabilities diffuse unusually fast. A new factory takes years to build and a new rail network decades. A new software capability delivered through the cloud can reach every customer of a provider the day it is released. When a competitor can rent the same computing infrastructure as a market leader and deploy a new application in months, the protective value of scale in physical assets falls. This compresses the time an incumbent has to respond, which is exactly the feature of extinction events that Siebel wants his readers to feel. The speed with which generative AI tools spread after 2022, reaching hundreds of millions of users within a couple of years, is a later illustration of the same property, though it arrived after his book was written. Finally, the metaphor captures something about interdependence. In an ecosystem, the loss of one species can cascade through food chains. In an economy, the digitization of one firm changes the expectations of its customers, suppliers and regulators, which then presses on everyone else. When one bank lets customers open an account from a phone in minutes, every bank is measured against that standard. When one manufacturer offers guaranteed uptime backed by sensor monitoring, rivals must match it. The environment changes not only because technology improves but because leading adopters reset what counts as normal. Where the analogy strains A good strategy paper uses a metaphor and then checks it. The extinction analogy has at least four weaknesses that a critical reader should name. The first is agency. Species do not choose to adapt; firms do. The whole point of Siebel's book is that executives can act, and that those who transform will survive. That is a very different claim from biological extinction, in which individual organisms have no strategic options. The metaphor therefore dramatizes the threat while obscuring the fact that the outcome depends on managerial decisions, investment and execution. Read strictly, the analogy even undermines the book's call to action, because it implies that fate is determined by traits already in place. The second is measurement. Turnover in lists such as the Fortune 500 is a blunt instrument. Companies leave the list because they are acquired, which may represent success for their shareholders; because they are taken private; because they are split up; or because other firms simply grow faster. A company can drop off the list while remaining healthy and profitable. Turnover in large-company rankings has also been high in earlier eras, including periods well before cloud computing. The statistic is suggestive, but it does not by itself show that digital technology is the cause, or that the current period is unusual. A student citing it should say so. The third is causal attribution. Many of the famous corporate collapses of recent decades had multiple causes: financial leverage, regulatory change, globalization, management error, and shifts in consumer taste. Assigning them all to a single technological shock makes for a clean narrative but weak analysis. Kodak, often cited as a victim of digital photography, in fact built one of the first digital cameras in its own laboratories in the 1970s; its difficulties had as much to do with the economics of its film business and its strategic choices as with any failure to see the technology coming. The fourth is interest. A mass-extinction narrative creates urgency, and urgency sells transformation projects. Consultants, platform vendors and cloud providers all benefit when executives believe that inaction is fatal. Siebel's company is one of those vendors. This does not make the narrative false, but it means a reader should look for independent evidence of the threat in any specific industry rather than accept the general claim. The upshot is that the metaphor is best treated as a hypothesis about environmental change, not a law. The useful question for any given firm is not whether a mass extinction is happening in general but whether the specific conditions that would make one likely are present in that firm's industry. Those conditions are the subject of the rest of this guide, and they are collected into a vulnerability rubric in the final chapter. Survivors, and what they suggest If the extinction metaphor were literally true, large incumbents would be doomed as a class. The record of the past two decades is more varied, and the variety is instructive. Several of the most successful digital transformations have been carried out by old, large companies rather than by start-ups. Microsoft is the most striking example. In the late 2000s it was widely portrayed as a company tied to a declining franchise in desktop software. Under Satya Nadella, who became chief executive in 2014, it reoriented the business around its Azure cloud platform and subscription software, and it became one of the three dominant cloud infrastructure providers. The case matters for Siebel's argument because it shows that an incumbent can reposition itself around exactly the technologies he identifies, and that doing so required a change in strategy and culture directed from the top, which is the leadership model he prescribes. Industrial companies offer a second kind of evidence. Farm equipment manufacturer John Deere has for years invested heavily in connected machinery, satellite positioning and data services that help farmers plan planting, spraying and harvesting. Its machines now generate streams of sensor data, and part of its value proposition has moved from steel to software. Whatever one thinks of the commercial and repair-rights controversies this has provoked, the case illustrates a firm treating the Internet of Things and analytics as a way to deepen, rather than lose, its position. Retail provides a third. Walmart, long the model of physical retail efficiency, faced intense pressure from Amazon and responded with large investments in e-commerce, online grocery ordering and store-based fulfilment. The outcome shows that incumbents can use existing assets, in this case a dense network of stores that double as distribution points, as advantages in a digital competition rather than as dead weight. These cases do not refute Siebel. They support his central prescription: transformation is possible, and it is led by executives who treat technology as a redesign of the business. What they complicate is the metaphor. In nature, the traits that fitted an animal to the old world are fixed. In business, a large installed base of customers, data and physical infrastructure can become either a burden or an advantage depending on how it is used. A student analyzing a particular firm should therefore resist treating size or age as a simple risk factor and ask instead whether those assets are being connected to the new capabilities. The survivors also point to a subtler lesson. In each case the firm did not adopt a single technology in isolation. Microsoft combined cloud infrastructure with data services and, later, artificial intelligence. Deere combined sensors, connectivity, cloud storage and predictive models. Walmart combined online ordering with inventory data and logistics optimization. The value came from the combination, which is precisely the convergence argument at the center of Siebel's book, and the reason this guide treats convergence as his most important idea. Turning the metaphor into an analytical question It helps to restate Siebel's thesis in terms that can be tested. Reduced to its logical core, it makes four claims. Firms face a technology shift that alters the cost and capability frontier across industries. The shift is large enough that marginal improvement will not keep a firm competitive. The traits that impede adaptation are mainly organizational and architectural, not a lack of access to the technology itself. And the firms that adapt earliest will capture disproportionate gains because of increasing returns. Each claim points to evidence a student can look for. Evidence for the first would be large, measurable differences in unit costs or cycle times between digitally mature and immature firms in the same industry. Evidence for the second would be cases in which incremental adopters lost ground despite investing. Evidence for the third would be firms with ample budgets and access to the same cloud providers that nonetheless failed to transform. Evidence for the fourth would be widening gaps in margins or market share over time. The mapping between the biological source and the business target can be set out compactly, and Table 1 does so, with a column noting where the analogy holds and where a careful analyst should qualify it. Table 1. The extinction metaphor mapped to business. Biological concept Business analog in Siebel Where it holds Where to qualify it Stable equilibrium Settled industry structure Long periods of stable leaders Stability varies widely by sector Environmental shock Convergence of four technologies Cost and capability shift is real Not the only cause of decline Mass extinction Wave of corporate failure High turnover among large firms Turnover includes mergers and growth elsewhere Fitness to old environment Legacy systems and silos Incumbency can slow change Incumbents also hold data and customers Adaptive radiation New digital leaders emerge Increasing returns reward early movers Firms, unlike species, can choose to adapt Framed this way, Siebel's opening chapter becomes the thesis statement of a research program rather than a prophecy. The rest of the book, and this guide, examines whether the machinery behind the thesis actually works. Key Takeaways · Siebel uses the evolutionary ideas of punctuated equilibrium and mass extinction to argue that a technology shock is resetting the conditions for corporate survival. · The metaphor implies speed, indiscriminate risk to incumbents, and large opportunities for firms that adapt. · It captures real economics: drastic cost shifts, increasing returns to data, and the drag of legacy organization. · It strains on agency, measurement, causal attribution, and the commercial interest of those who promote it. · The most useful move is to restate the thesis as testable claims and look for industry-specific evidence. Review Questions 1. What three implications does the mass extinction metaphor carry for incumbent firms, and which is most important to Siebel's argument? 1. Why is turnover on the Fortune 500 list an imperfect measure of technology-driven corporate failure? 2. How does Siebel's account of incumbent vulnerability differ from Christensen's model of disruptive innovation? 3. Explain how the element of managerial agency both weakens the biological analogy and supports Siebel's call to action. 4. Choose an industry and state what evidence would confirm or refute the claim that it is undergoing a "mass extinction." Chapter 2: Earlier Waves of Information Technology Siebel's claim that the present moment is a mass extinction depends on a comparison. If every earlier wave of computing had also been described as world-ending, and if most firms had adapted without trouble, the claim would look like the familiar enthusiasm of a technology salesman. So Siebel places digital transformation in a longer history of enterprise computing, one he watched from the inside for much of its course. His purpose is to show two things: that each wave changed how businesses operate, and that the present wave differs from its predecessors in scope and speed. This chapter reconstructs that history, explains the pattern Siebel draws from it, and evaluates whether the pattern supports his conclusion. From mainframes to the internet The first commercial computers of the 1950s and 1960s were mainframes: large, expensive machines housed in dedicated rooms and operated by specialists. IBM dominated this era, and its System/360 family, announced in 1964, established the idea of a compatible line of machines on which a company could run its accounting, payroll and inventory applications. The mainframe era automated back-office record keeping. It replaced clerks and ledgers with batch processing, in which transactions were collected and run through the computer at intervals, often overnight. Minicomputers, led by Digital Equipment Corporation from the mid-1960s, brought computing down in size and price, allowing departments, laboratories and factories to own machines rather than share a central one. Personal computers arrived in force after the IBM PC of 1981, putting computing on individual desks and creating a mass market for spreadsheets, word processing and, later, networked office work. Each of these shifts dispersed computing power further from the center of the organization. The next wave, and the one Siebel knows most intimately, was client-server computing and the rise of packaged enterprise applications in the late 1980s and 1990s. Instead of writing custom programs for every business function, companies bought standardized software suites. Enterprise resource planning (ERP) systems, of which SAP's R/3, released in 1992, became the leading example, integrated finance, manufacturing, procurement and human resources. Customer relationship management systems, the category Siebel Systems led, did the same for sales, marketing and customer service. These systems promised a single, consistent record of a company's operations and processes, and they encouraged firms to redesign their workflows around best practices built into the software. Then came the internet. The commercial web, symbolized by the public offering of Netscape in 1995, connected companies directly to customers and suppliers. Electronic commerce, online banking and web-based customer service changed the interface between firms and markets. The internet boom and its collapse in 2000 and 2001 are a reminder that transformative technologies can be both real and overhyped at the same time: many of the business models of that era failed, yet the underlying shift to online commerce proved permanent. After the internet came two further developments that set the stage for Siebel's four technologies. The first was software as a service, in which applications are delivered over the internet and paid for by subscription rather than installed on a company's own servers. Salesforce, founded in 1999, built its business on this model and competed directly with the kind of installed CRM software Siebel Systems sold. The second was mobile computing, accelerated by Apple's iPhone in 2007, which put a connected computer in the pockets of billions of people and made constant data generation a normal part of daily life. The pattern Siebel draws From this history Siebel extracts a pattern with several features. The most important is acceleration. Each wave arrived faster than the last and diffused more quickly. Mainframes took decades to spread; the internet reached mass adoption within about a decade; mobile applications spread within a few years. Behind this acceleration lies the steady improvement in computing performance and cost associated with Moore's law, Gordon Moore's 1965 observation that the number of components on an integrated circuit was doubling at a regular interval. For half a century, the cost of a given amount of computation fell dramatically, which repeatedly made new applications economical. A second feature is expanding scope. Early computing automated narrow tasks in the back office. Enterprise applications integrated whole functions. The internet connected the firm to the outside world. Each wave touched more of the organization than the one before. In Siebel's telling, the present wave is the first that reaches every function simultaneously, from engineering and operations to sales, finance and human resources, because it is not a new kind of application but a new capability that can be applied to any process that generates data. A third feature is winner turnover among vendors. Each wave produced new leading technology companies and humbled some old ones. The mainframe leaders did not dominate personal computing; the packaged-software leaders did not dominate the internet; many internet-era leaders were displaced by the cloud and mobile giants. Siebel has lived this pattern personally, having built a company that led one wave and was absorbed into Oracle as the next wave, cloud-delivered software, gained momentum. A fourth feature, which is more implicit, is that each wave shifted the locus of competitive advantage. When computing was scarce, owning it was an advantage. When software was packaged, implementing it well was an advantage. When everyone could buy the same software, advantage moved to how a company used information and to its relationships with customers. Siebel's argument is that in the present wave, advantage moves to the ability to integrate data across the whole enterprise and apply predictive models to it. Why Siebel thinks this wave is different The comparison with earlier waves allows Siebel to make his key distinction. Earlier waves, in his account, mainly digitized existing processes. A company that installed ERP software still ran the same basic business, only more efficiently and with better records. The internet let firms sell through a new channel, but the product and the operating model often stayed recognizable. The present wave, he argues, changes what the enterprise can know and predict. When every asset is instrumented with sensors, when all transactions and interactions are captured and stored at low cost, and when machine learning models can find patterns across that data, the firm can move from reacting to events to anticipating them. A utility can predict which transformer will fail. A bank can predict which transaction is fraudulent before it clears. A manufacturer can predict demand at a granular level and adjust production. Siebel treats this move from descriptive to predictive operation as a qualitative change, not merely a faster version of what came before. Earlier waves can be compared with the present one on a few consistent dimensions, as Table 2 shows. The dates are approximate and the characterizations are the guide's summary, intended to make Siebel's argument about scope and speed easy to see. Table 2. Waves of enterprise computing compared. Wave Approximate era Main change Scope in the firm Typical leaders Mainframe 1950s to 1970s Batch automation of records Back office IBM Minicomputer and PC 1965 to 1990s Distributed computing Departments and desks DEC, IBM, Microsoft Client-server applications Late 1980s to 2000s Integrated process software Whole functions SAP, Oracle, Siebel Systems Internet and SaaS Mid-1990s to 2010s Online channels and subscription software Firm-to-market interface Amazon, Salesforce Convergence of four technologies 2010s onward Prediction across all operations Entire enterprise Cloud providers and AI platforms Testing the historical argument The historical account is persuasive in outline, but a critical reader should test it at three points. The first is the claim of uniqueness. Every wave was described by contemporaries as unprecedented. Writers in the 1990s argued that the internet would abolish intermediaries and transform every industry, and some of that proved true while much did not. The fact that Siebel's argument sounds like the claims made at the start of earlier waves does not make it wrong, but it should make a student cautious about accepting "this time is different" without specific evidence. The strongest version of Siebel's case rests not on rhetoric but on a mechanism: prediction at scale changes decisions that previously could not be made at all. That mechanism can be checked in particular industries. The second is the pace of adoption. Siebel emphasizes acceleration, and the diffusion of consumer technologies has indeed quickened. Enterprise adoption is slower and more uneven. Large organizations carry legacy systems, regulatory constraints and skill shortages that delay change. Many surveys of enterprise AI in the years after the book's publication found that a large share of pilot projects never reached full production. Siebel himself acknowledges this problem, and in fact builds his commercial argument on it: he says many companies fail because they try to assemble their own technology stacks. But it tempers the idea that the transformation is sweeping through all firms at once. The third is the productivity puzzle. Economists have long noted that major technologies often take years or decades to show up in aggregate productivity statistics. Robert Solow observed in 1987 that computers could be seen everywhere except in the productivity figures. Later research by economists such as Erik Brynjolfsson argued that general-purpose technologies require large complementary investments in organization, skills and processes before their benefits appear, a pattern sometimes described as a productivity J-curve. This literature is useful for evaluating Siebel. It supports his insistence that transformation is organizational, not merely technical. It also suggests that the gains from convergence may arrive more slowly and more unevenly than the extinction metaphor implies. The sediment of earlier waves One implication of Siebel's history deserves separate treatment because it shapes almost every real transformation effort. Computing waves do not replace one another cleanly. They leave sediment. A large bank, insurer or government agency today typically runs mainframe systems written decades ago alongside client-server applications from the 1990s, web front ends from the 2000s, and cloud services added in the last few years. Each layer holds part of the organization's data, in its own format, governed by its own rules. The practical cost of this sediment became visible in the spring of 2020. As unemployment claims surged at the start of the COVID-19 pandemic, several American states found that their benefit systems ran on old mainframe software written in COBOL, a programming language dating from 1959, and struggled to modify them quickly. The governor of New Jersey publicly appealed for volunteers with COBOL skills. The episode is a vivid illustration of a general problem: the older layers of an enterprise's technology often encode its most essential processes, and they are the hardest to change. This sediment is usually described as technical debt, a term for the accumulated cost of past shortcuts and outdated systems that must eventually be repaid through rework. For Siebel's argument, technical debt matters in two ways. It explains why incumbents struggle to move at the speed of new entrants, which have no legacy layers to integrate. And it explains why he places such weight on data integration, the subject of a later chapter. An enterprise cannot make predictions across its operations if the data describing those operations is scattered across systems from five different eras that were never designed to talk to one another. It also suggests a refinement of the extinction metaphor. The obstacle facing incumbents is not simply that they lack new technology, which can be rented from cloud providers by anyone. It is that their existing technology, and the processes built around it, must be connected to the new capabilities or retired. The cost and risk of doing so varies enormously between firms, and it is one of the most useful variables a student can examine when assessing a company's exposure to disruption. Two firms in the same industry, with similar revenues and similar ambitions, can face very different transformation bills because one modernized its core systems a decade ago and the other did not. There is a counterpoint worth noting. Legacy systems are often old because they work. Mainframes still process a large share of the world's card transactions and airline bookings because they are reliable and secure at very high volumes. A transformation plan that treats every old system as a liability can destroy value. The better approach, and one consistent with Siebel's architecture, is to leave stable systems of record in place where they perform well and to build a layer above them that extracts, integrates and analyzes their data. Whether that layer is best bought or built is a question this guide returns to in detail. What the history contributes to strategy analysis For a student, Siebel's history offers a practical lens. When analyzing a firm, ask which waves it has fully absorbed and which it has not. A company still running core operations on decades-old systems, with data trapped in separate applications that do not communicate, has not completed the client-server integration that earlier waves promised, let alone the convergence Siebel describes. Its path to prediction is longer, because it must first solve integration problems left over from earlier eras. The history also warns against assuming that today's technology leaders are permanent. Siebel's own career shows how quickly a dominant vendor can be overtaken. That lesson applies to the platform companies of the present, including hyperscale cloud providers and AI software firms, and a strategy paper that recommends committing to a particular vendor should consider the risk that the vendor itself may be displaced by the next wave. The arrival of generative AI after 2022, which shifted attention and investment toward a new set of model developers, is a live example of that risk. Finally, the history clarifies what exactly is new. The novelty is not computers, data or even statistics, all of which businesses have used for decades. It is the combination of practically unlimited computing capacity, comprehensive data capture, learning algorithms that improve with data, and a physical world increasingly fitted with sensors. The following four chapters examine each of these elements in turn, beginning with the one that made the others economically practical: elastic cloud computing. Key Takeaways · Siebel places digital transformation in a sequence of computing waves from mainframes through client-server, the internet, software as a service and mobile. · He identifies a pattern of accelerating diffusion, expanding scope within the firm, turnover among technology leaders, and shifting sources of advantage. · His central distinction is that earlier waves digitized existing processes while the current wave enables prediction across the whole enterprise. · Claims that "this time is different" recur in every wave and should be tested against specific mechanisms and evidence. · Economic research on general-purpose technologies supports Siebel's emphasis on organizational change but suggests gains may arrive slowly. Review Questions 1. What four features does Siebel's history of computing waves reveal, and which does he rely on most heavily? 1. How does the shift from digitizing processes to predicting outcomes support the claim that the current wave is qualitatively different? 2. Why should a strategy analyst be cautious about claims of technological uniqueness, and what evidence would make such a claim credible? 3. How does the productivity J-curve argument both support and challenge Siebel's thesis? 4. What does Siebel's own career suggest about the durability of technology vendors, including his own company? Chapter 3: Elastic Cloud Computing Of Siebel's four technologies, cloud computing is the foundation. Without it, the other three would remain expensive experiments available only to the largest firms. Big data requires cheap, effectively unlimited storage. Artificial intelligence requires bursts of intense computation for training models. The Internet of Things requires infrastructure that can absorb streams of readings from millions of devices. Siebel's argument is that elastic cloud computing supplies all three at a price and on a timescale that no company's own data center could match. This chapter explains what the cloud is, why its elasticity matters economically, how the market has developed since the book appeared, and where Siebel's enthusiasm needs qualification. What elasticity means The most widely cited definition of cloud computing comes from the United States National Institute of Standards and Technology, whose 2011 publication identified five essential characteristics. Customers can obtain computing resources on demand without human interaction with the provider. Services are accessible over the network from many kinds of device. The provider pools resources to serve many customers at once. Capacity can be expanded and released rapidly, in some cases automatically. And usage is measured, so customers pay for what they consume. The same document distinguished three service models, which remain standard vocabulary. Infrastructure as a service (IaaS) provides raw computing, storage and networking. Platform as a service (PaaS) adds managed tools on which developers build applications. Software as a service (SaaS) delivers finished applications. Siebel places the emphasis on the fourth characteristic, which gives his chapter its adjective. Elasticity is the ability to scale capacity up and down in step with demand. It matters because computing demand in a modern enterprise is highly uneven. A retailer needs far more capacity during holiday sales than in February. A bank's risk calculations spike at the end of a quarter. Training a large machine learning model may require thousands of processors for days and then none at all. In a traditional data center, a company had to buy enough hardware for its peak load and leave much of it idle the rest of the time. In an elastic cloud, it rents capacity for the hours it needs and releases it afterward. The economic consequence is a shift from capital expenditure to operating expenditure. Instead of committing millions of dollars up front to servers that will depreciate over several years, a firm pays for computing as a variable cost. This lowers the barrier to experimentation. A team can test an idea with a few hundred dollars of computing, and if it works, scale it up without a procurement cycle. Siebel regards this as one of the main reasons the pace of innovation has increased: the cost of trying something new has collapsed. A second consequence is access to scale that no single company could build. The largest cloud providers operate global networks of data centers, buy hardware in enormous volumes, design their own chips, and spread fixed costs over millions of customers. A mid-sized manufacturer renting capacity from them gains, in effect, the infrastructure of a technology giant. For Siebel, this democratization is central. It means that the ability to run advanced AI is no longer confined to a handful of internet companies, which is why he argues the transformation will reach every industry. A third consequence is speed of provisioning. Before the cloud, obtaining a new server in a large company could take weeks or months of approvals, purchasing and installation. In the cloud it takes minutes. That change, multiplied across thousands of projects, alters how quickly an organization can move, and it is one reason Siebel insists that transformation projects should be measured in months rather than years. Siebel's position on the cloud Siebel treats the move to the public cloud as inevitable and largely complete in direction, if not in execution. He describes the major providers, Amazon Web Services, Microsoft Azure and Google Cloud, as having built capabilities that corporate data centers cannot match on cost, security or reliability, and he expects enterprise computing to migrate to them on a large scale. He also recognizes that some organizations, particularly governments and firms in regulated industries, will retain private infrastructure for some workloads, and that many will operate in hybrid arrangements that combine the two. An important strand of his argument concerns independence from any single provider. Siebel's company designs its software to run on several clouds, and the book presents this as an advantage for customers, who can avoid being locked into one vendor's proprietary services. This is a reasonable strategic point, and it is also a feature of his company's product. A student should recognize both. Multi-cloud portability does reduce dependence on one provider. It also adds a layer of software, and a vendor relationship, between the customer and the cloud. Siebel also stresses the security argument. Large cloud providers spend heavily on physical and cyber security and employ specialists that most companies cannot afford. In his view, a well-managed public cloud is often more secure than a typical corporate data center. This view has become mainstream, but it comes with the qualification that security in the cloud is a shared responsibility: the provider secures the infrastructure, while the customer remains responsible for how it configures services, manages access and protects its data. Many publicized cloud breaches have resulted from customer misconfiguration rather than failures of the provider's own systems. The cloud market since the book Siebel wrote at a time when cloud adoption was growing rapidly but was still a minority of enterprise computing. Since then, the market has grown very large and has become highly concentrated. According to Synergy Research Group, which tracks spending on cloud infrastructure services, enterprise spending reached about 107 billion dollars in the third quarter of 2025 and about 129 billion dollars in the first quarter of 2026, with year-on-year growth of 35 percent in the latter period, the fastest rate the firm had recorded since 2021. Three providers dominate. Their shares in those two quarters are set out in Table 3. Table 3. Shares of enterprise spending on cloud infrastructure services. Provider Q3 2025 share Q1 2026 share Amazon Web Services 29% 28% Microsoft Azure 20% 21% Google Cloud 13% 14% All others combined 38% 37% Source: Synergy Research Group, market share releases of November 2025 and April 2026. Three developments stand out for a student applying Siebel's framework in 2026. The first is concentration. Three firms account for roughly 63 percent of the market. Siebel's argument that the cloud democratizes computing is true in the sense that any company can rent capacity. It is also true that the supply of that capacity is controlled by a small number of very large firms. This raises strategic questions about bargaining power, pricing and dependence that the book touches only lightly. Regulators have noticed. The United Kingdom's Competition and Markets Authority investigated the cloud market, and the European Union's Data Act, which began to apply in September 2025, includes provisions intended to make it easier for customers to switch between cloud providers. In 2024 the three leading providers each announced that they would waive data transfer fees for customers moving their data out, a change widely seen as a response to regulatory pressure. The second is generative AI as a driver of cloud demand. Synergy and other analysts have attributed much of the recent acceleration in cloud growth to the computing needs of generative AI, both for training large models and for running them at scale. This development validates Siebel's claim that AI and cloud are tightly linked, and it has changed the cost structure of the market. Specialized processors, particularly graphics processing units, have been scarce and expensive, and providers have committed enormous capital to new data centers and the electricity to run them. The third is the rise of specialized providers, often called neoclouds, which focus on renting AI computing capacity. Synergy's April 2026 report named companies such as CoreWeave, Nebius and Crusoe among the fast-growing suppliers, noting that several now rank among the thirty largest cloud providers and hold a much larger share of AI-specific workloads than of the market as a whole. Oracle has also grown its cloud infrastructure business rapidly on the strength of AI demand. The picture is therefore not simply one of three permanent giants. A new segment has opened, driven by the very technology Siebel placed at the center of his argument, and it is a reminder of the vendor turnover discussed in the previous chapter. A worked illustration of elastic economics The logic of elasticity is easiest to grasp through a simplified example. The numbers below are hypothetical round figures chosen for clarity, not market prices. Suppose a manufacturer wants to train a model that predicts equipment failures across its plants. Training the model requires the equivalent of 1,000 processor-days of computation, and the data science team expects to retrain it once a month as new sensor data arrives. Between training runs, the model needs only modest capacity to score new readings. Under the traditional approach, the firm would size its own cluster to finish a training run in a reasonable time, say ten days, which means buying about 100 processors. For roughly two-thirds of every month, most of those processors sit idle. The firm also carries the cost of the building space, power, cooling and staff to run them, and it must decide on the purchase before it knows whether the model will work at all. Under the elastic approach, the firm rents 1,000 processors for a single day, completes the training run overnight, and releases them. It pays for the same total amount of computation, but it receives the result nine days sooner, holds no idle hardware, and commits nothing up front. If the project fails, the firm has spent the cost of a few runs rather than the cost of a cluster. If the project succeeds and the firm wants to train models for every plant, it rents more capacity for the same short window. Three lessons follow. The first is that elasticity converts time into money in a new way: a firm can buy speed by renting more capacity for a shorter period at roughly the same total cost, which was impossible when capacity was fixed. The second is that elasticity changes the economics of failure. When experiments are cheap and reversible, an organization can afford to try many and keep the few that work, which is how most successful AI programs actually develop. The third is that the advantage depends on utilization. If the manufacturer's cluster were busy every hour of every day with other work, owning it might well be cheaper, which is why the cost argument must always be tested against the actual pattern of demand. The example also shows why Siebel ties cloud computing so tightly to AI. Machine learning is unusually bursty, unusually experimental and unusually hungry for computation. It is close to the ideal workload for elastic infrastructure, and it is hard to imagine enterprise AI reaching its present scale without it. The limits of the cloud argument Siebel's case for the cloud is strong, but it is not unconditional, and a rigorous analysis should weigh several counterarguments. Cost at steady state. Elasticity is most valuable for variable workloads. For large, predictable workloads that run continuously, owning hardware can be cheaper than renting it. Some companies have publicly moved workloads off the public cloud for this reason. The software company 37signals, maker of the Basecamp project management tool, announced in 2022 that it would leave the public cloud for its own hardware, citing cost, and its executives have since written about the savings they expect. Most analysts regard such repatriation as a niche trend rather than a reversal, but it demonstrates that the economic argument depends on the shape of demand. Cost control. Pay-as-you-go pricing makes costs easy to incur and hard to predict. Many organizations have discovered cloud bills far higher than planned because teams provisioned resources and forgot to release them, or because data transfer and storage charges accumulated. A discipline known as FinOps, which brings financial accountability to cloud spending, has emerged in response. In a strategy paper, the cloud should be treated as a change in cost structure that requires new management controls, not as an automatic saving. Concentration risk. When many companies depend on the same provider, a single failure can have wide effects. On 20 October 2025, a major outage in Amazon Web Services' US-East-1 region, traced to a problem with the domain name system resolution of its DynamoDB database service, disrupted a large number of widely used applications and websites for hours. Events like this do not negate the reliability advantages of the cloud, since most corporate data centers also fail, but they show that concentrated infrastructure creates correlated risk. A firm's resilience plan should consider whether it can operate across regions or providers. Lock-in. Cloud providers offer many proprietary services that are convenient but hard to replicate elsewhere. The more a company uses them, the more costly it becomes to switch. This is the concern Siebel's multi-cloud argument addresses, though, as noted, adding an independent platform layer substitutes one form of dependence for another. The real choice is not between dependence and independence but between different distributions of dependence. Sovereignty and regulation. Data protection rules, national security concerns and industry regulation limit where some data may be stored and processed. European governments in particular have pursued forms of "sovereign cloud" to keep sensitive data under local jurisdiction. For multinational firms, the cloud decision is increasingly shaped by law as well as economics. Energy and physical constraints. The expansion of AI-driven cloud capacity has made electricity supply, cooling and grid connection significant constraints on growth. This was not a central concern when Siebel wrote. It now affects the location, cost and speed of new capacity, and it links the cloud to the energy transition and to public policy on power generation. Using the cloud in a convergence analysis For the purposes of this guide's convergence matrix, cloud adoption is best assessed not by asking whether a firm "uses the cloud," which almost every firm now does, but by asking three sharper questions. What share of its critical workloads, and especially its data and analytics workloads, runs on elastic infrastructure? Can it scale computing quickly for new analytical or AI projects without a lengthy procurement process? And does it have the financial and security controls to use the cloud without runaway costs or exposure? A firm that answers well on all three has the foundation Siebel describes. A firm that has moved email and office applications to the cloud but keeps its operational data in on-premises systems it cannot easily connect has not yet laid that foundation, whatever its public statements claim. The distinction matters because, as the next chapter shows, the real bottleneck in Siebel's model is not computing power but data. Key Takeaways · Siebel treats elastic cloud computing as the foundation that makes big data, AI and IoT economically practical for ordinary firms. · Elasticity converts computing from a capital expense into a variable operating cost, lowers the cost of experiments, and speeds provisioning. · The cloud market has grown rapidly and is concentrated: three providers held about 63 percent of spending in late 2025, according to Synergy Research Group. · Generative AI has accelerated cloud demand and given rise to specialized AI-focused providers. · Counterarguments include steady-state cost, weak cost control, concentration risk, lock-in, regulation and energy constraints. Review Questions 1. Explain the economic significance of elasticity and why it matters particularly for training AI models. 1. What is the shared responsibility model of cloud security, and how does it qualify Siebel's claim that the cloud is more secure? 2. How does market concentration among cloud providers complicate the argument that the cloud democratizes computing? 3. Under what conditions might owning computing infrastructure be more economical than renting it? 4. Why does multi-cloud portability redistribute, rather than eliminate, a firm's dependence on vendors? Hashtags: #TheConvergenceMatrix #DigitalTransformation #ThomasSiebel #TechnologicalConvergence #CloudComputing #BigData #ArtificialIntelligence #InternetOfThings #EnterpriseAI #PunctuatedEquilibrium #MassExtinctionThesis #DigitalDisruption #LegacySystems #TechnicalDebt #DataIntegration #ModelDrivenPlatforms #PredictiveOperations #ElasticCloudComputing #IncreasingReturns #GeneralPurposeTechnology #ProductivityJCurve #BuildVersusBuy #CEOLedTransformation #StrategicVulnerability #FutureOfDigitalTransformation
- The Dietetic Blueprint (A Student's Guide to Krause and Mahan's Food & the Nutrition Care Process)
Download the Book (PDF): Introduction This is the most trusted, widely used textbook in dietetics education worldwide. It is the ultimate authority on the Nutrition Care Process (NCP). However, navigating its massive chapters on lifecycle nutrition, pediatric specialties, and complex medical nutrition therapies can easily become a blur of diagnostic codes and physiological charts for students trying to cram for a clinical evaluation. This book structures the clinical data into a predictable diagnostic model. It is explicitly written to explain Krause and Mahan's Food & the Nutrition Care Process, organizing the exhaustive textbook into the strict ADIME framework (Assessment, Diagnosis, Intervention, Monitoring, and Evaluation). We break down exactly how to translate physiological symptoms into standardized nutrition diagnoses. Packed with clear case study rubrics, this guide is your key to mastering clinical dietetic practice. This is an educational study aid for students, not clinical advice: every value, equation, threshold and management approach described here must be checked against your own textbook, your institution's protocols and the current primary guidelines before it is used in the care of any real patient. The book this guide explains Krause and Mahan's Food & the Nutrition Care Process has been the spine of dietetics education in the English-speaking world for more than seventy years. It began as Food, Nutrition and Diet Therapy under Marie V. Krause, and its authorship passed through L. Kathleen Mahan and Sylvia Escott-Stump before arriving with the present editors, Janice L. Raymond and Kelly Morrow. The current edition at the time of writing is the sixteenth, published by Elsevier in 2022 under the editorship of Raymond and Morrow; the fifteenth edition, from 2020 to 2021, remains in wide circulation in programmes that have not yet adopted the newer printing, and where the two differ in structure this guide notes it. The chapter-level organisation described here follows the sixteenth edition. If your programme has issued the fifteenth, almost everything below still applies, because what changed between the editions is the currency of the evidence rather than the architecture of the reasoning. The textbook's authority rests on two things that are easy to take for granted. The first is scope: it moves from cellular nutrient metabolism through the whole human lifespan and then across essentially every organ system that nutrition can affect, which is why it is often the only clinical nutrition text a student is required to buy. The second is that it was among the first major texts to reorganise itself around the Nutrition Care Process rather than around disease categories. That decision is the reason the book is difficult in exactly the way students complain about, and it is also the reason the book is worth the difficulty. A text organised by disease teaches you a list. A text organised by process teaches you a method that works on a disease you have never seen. The problem this guide solves The failure mode in clinical nutrition coursework is not laziness. It is that students read the textbook the way it is printed — front to back, chapter by chapter, absorbing lists of laboratory values, nutrient functions and dietary modifications — and then arrive at a case study or a simulated patient encounter and find that none of it comes out in a usable order. They can recite the sodium restriction for heart failure and the protein recommendation for haemodialysis, but they cannot look at a set of notes and say what the nutrition problem actually is, what is causing it, and what evidence in front of them proves it. That gap exists because the textbook's clinical chapters are reference material, and reference material is organised for lookup, not for reasoning. The reasoning lives in the earlier chapters on the Nutrition Care Process, and most students skim them in week two and never return. The controlling idea of this guide is straightforward: the Nutrition Care Process is not an introductory topic in clinical dietetics but the operating system on which the rest of the textbook runs, and once you can execute it reliably, every clinical chapter in Krause and Mahan becomes a set of parameters you drop into a structure you already know. Assessment tells you what to collect and what to compare it against. Diagnosis forces you to name one problem in a form that commits you to a cause and to evidence. Intervention obliges you to aim at the cause you named. Monitoring and evaluation makes you say in advance what would count as improvement. The disease chapters supply the specifics — this population's comparative standards, this condition's characteristic deficits, this guideline's protein target — but they never supply the structure, and a student who has the specifics without the structure produces care plans that are collections of good intentions rather than clinical reasoning. How this guide is organised The guide works outwards from that structure. It begins with the Nutrition Care Process itself and its model, sets out the four steps as the Academy of Nutrition and Dietetics defines them, and shows how the ADIME note — the documentation format you will be graded on — is simply the four steps written down. From there it spends three chapters on assessment, because assessment is where most student errors originate and where the textbook's sheer density does the most damage: one chapter on the five categories of assessment data and how to gather them, one on the quantitative work of estimating energy, protein and fluid needs and of reading anthropometric and growth data, and a treatment of malnutrition identification that keeps the competing consensus frameworks distinct rather than blurring them together. The diagnosis chapter is the hinge of the book. It sets out the standardised diagnostic terminology and its domains, works through the PES statement in detail, catalogues the errors that cost marks and, more importantly, cost patients good care, and offers a quality rubric you can apply to your own statements before anyone else does. The intervention and monitoring chapters then follow the logic forward into the nutrition prescription, the delivery of food and nutrients, education and counselling — including the behaviour change theory that underpins them — and into the discipline of naming indicators and criteria before you measure anything. The second half applies all of this. Life-cycle nutrition comes first, because pregnancy, infancy, childhood and old age each change the comparative standards rather than the method. Then three chapters of medical nutrition therapy, grouped so that conditions sharing a physiological mechanism sit together: metabolic and cardiovascular disease; gastrointestinal, hepatobiliary, pancreatic and renal disease; and the cluster of cancer, critical illness, food allergy and nutrition support, where enteral and parenteral feeding and the risk of refeeding syndrome are treated in full. The final chapter puts everything to work on three fully developed fictional cases with complete ADIME write-ups, together with a grading rubric devised for this guide, so that you can see what a strong answer looks like and mark your own against it. Throughout, the aim is to be faithful to the source. Where the guide describes what Raymond, Morrow and their contributors argue, it says so. Where it adds its own framing, a rubric, a mnemonic or a critique of how a framework performs in practice, it says that too. And where a clinical threshold matters, it is attributed to the primary guideline it comes from — the Academy's own terminology, the Global Leadership Initiative on Malnutrition, the American Society for Parenteral and Enteral Nutrition, the Kidney Disease Outcomes Quality Initiative, the American Diabetes Association — rather than to the textbook alone, because guidelines are revised on their own schedules and you will need to track them long after this edition of Krause and Mahan is superseded. Chapter 1: The Nutrition Care Process and Model Before the Nutrition Care Process existed, two competent dietitians could see the same patient, agree entirely about what was wrong, and write two notes that shared almost no vocabulary. One would record "poor intake secondary to nausea"; the other "patient not meeting estimated needs." Neither was wrong. Neither could be counted. If you wanted to know how many patients in a hospital had a nutrition problem that dietetic intervention had resolved, there was no way to find out, because the profession had no standard way of saying what the problem was. The Academy of Nutrition and Dietetics adopted the Nutrition Care Process in 2003 to close that gap, and Krause and Mahan's Food & the Nutrition Care Process reorganised itself around it. Raymond and Morrow treat the process not as one topic among many but as the framework the clinical half of the book is built to serve. Understanding why it was adopted tells you what it is for, and what it is for tells you how to use it. What the process is, and what it is not The Nutrition Care Process is a systematic problem-solving method that registered dietitian nutritionists use to think critically about nutrition care and to make decisions that address nutrition-related problems. It is deliberately modelled on the problem-solving cycles used elsewhere in health care — nursing has one, physical therapy has one — because a profession that wants to be treated as a clinical discipline needs to be able to show that its decisions follow a defensible sequence rather than personal habit. Two things it is not. It is not a standard of care, in the sense of dictating what any given patient should receive; two dietitians using the process properly may reach different reasonable plans for the same patient. And it is not a documentation format, although a documentation format falls out of it. The process is the reasoning; ADIME is the record. The Academy's own description sets out four steps. In nutrition assessment, the practitioner collects and documents information across five categories: food or nutrition-related history; biochemical data, medical tests and procedures; anthropometric measurements; nutrition-focused physical findings; and client history. In nutrition diagnosis, the practitioner uses those data to identify and name the specific problem. In nutrition intervention, the practitioner selects an intervention directed at the root cause, or etiology, of that problem and aimed at relieving its signs and symptoms. In nutrition monitoring and evaluation, the practitioner determines whether the patient has achieved, or is progressing towards, the planned goals. Read that sequence closely and you can see that each step constrains the next. The diagnosis may only name problems the assessment has evidence for. The intervention must aim at the etiology the diagnosis named. The monitoring must measure the signs and symptoms the diagnosis cited. A student who understands that chain of constraint has understood most of what the process is trying to teach, and will never again write a care plan whose parts do not connect. The model around the process The four steps sit at the centre of a graphical scheme the Academy calls the Nutrition Care Process Model, and the model is worth more attention than students usually give it, because it encodes several assumptions that turn up in examination questions. The first is that the process is cyclical rather than linear. The Academy is explicit that practitioners frequently return to earlier steps as new information surfaces during an encounter — a patient mentions, twenty minutes in, that she stopped taking her pancreatic enzymes three weeks ago, and the assessment reopens, the diagnosis changes, and the intervention you had half drafted is abandoned. Students often imagine that the four steps are performed once, in order, and then filed. In practice the loop is tight and frequently re-entered, and monitoring and evaluation feeds directly back into reassessment. The second is that the process only operates inside a relationship. The model places the practitioner and the patient, client, group or population at the centre, and the whole apparatus depends on that relationship being functional. This is not decoration. A great deal of nutrition assessment data is self-reported, and self-reported data is only as good as the trust between the two people in the room. The third is that the practitioner brings something to the process that the process does not supply. The model surrounds the four steps with the strengths and abilities the dietitian brings: dietetics knowledge, evidence-based practice, critical thinking, collaboration, communication, and adherence to a code of ethics. The process is a method, not an autopilot. Two students given identical data will produce different diagnoses, and the difference will come from these competencies rather than from the framework. The fourth is that the process operates within an environment it does not control. The outer ring of the model represents the health care system, economic considerations, social systems and practice settings — the things that determine whether the intervention you designed is actually available to this patient. A textbook-perfect plan for an oral nutrition supplement three times daily is worthless if the patient cannot afford it and the facility does not stock it. Raymond and Morrow return to this point repeatedly in the clinical chapters, and it is the single most common thing that separates an academically correct student care plan from a clinically useful one. Finally, the model incorporates screening and referral at the entry point, and outcomes management at the exit. Screening is not part of the Nutrition Care Process — it is what identifies who enters it, and it is frequently performed by nurses or other staff rather than dietitians. Outcomes management is what happens to the data afterwards: aggregated across patients, the standardised terminology allows a department to demonstrate what its interventions achieve, which is the point of standardising the language in the first place. Standardised terminology: NCPT and the eNCPT A process that produces incomparable notes solves nothing, so the Academy developed a controlled vocabulary to go with it. Originally published as the International Dietetics and Nutrition Terminology, it is now the Nutrition Care Process Terminology, maintained in an online reference known as the eNCPT and revised on a regular cycle; the 2023 edition is current at the time of writing. The terminology is organised by process step and then by domain within each step. Assessment terms carry codes for their domains — food/nutrition-related history, anthropometric measurements, biochemical data and medical tests and procedures, physical exam findings, client history, and several supporting categories including comparative standards and assessment tools. Diagnosis terms fall into intake, clinical and behavioural-environmental domains, plus a nutrition situation category. Intervention terms cover goal and prescription planning and then implementation through food and nutrient delivery, education, counselling, coordination of care, population-based action and the context of the encounter. Monitoring and evaluation reuses most of the assessment domains, because the whole point is to remeasure what you measured before. Each term in the reference carries a definition, the data or indicators associated with it, and guidance on appropriate use. This matters practically: students routinely invent diagnoses that sound plausible but do not exist in the terminology, and a diagnosis that does not exist in the terminology cannot be aggregated, cannot be audited, and in most teaching hospitals will be returned for correction. A caution about the terminology that the textbook raises and that is worth taking seriously. Standardised language solves a communication problem and creates a risk. The risk is that a student learns to pattern-match a patient to a term rather than to understand the patient — reaching for "inadequate energy intake" because the number is low, without asking why it is low or whether the low number is even the problem worth naming. The terminology is a way of recording a judgement. It is not a substitute for making one. ADIME: the process written down ADIME is the acronym for the documentation format that mirrors the four steps: Assessment, Diagnosis, Intervention, Monitoring and Evaluation. It supplanted the older SOAP note in most dietetic practice because SOAP — subjective, objective, assessment, plan — has no place to put a nutrition diagnosis, and burying a diagnosis inside an assessment section defeats the purpose of having named it. The structure of an ADIME note is the structure of the process, and Table 1 sets out what belongs in each section and what does not. The most common structural error students make is to put interventions in the assessment section, usually in the form of a sentence like "will recommend high-protein diet" written before the diagnosis has been stated. That ordering error is not cosmetic. It shows that the writer decided on the plan before naming the problem, which is exactly the habit the process exists to break. Table 1. ADIME note structure mapped to the Nutrition Care Process. Section NCP step What belongs here Common misplacement A: Assessment Step 1 Data across the five categories, estimated needs, comparative standards, interpretation Planned interventions; the diagnosis itself D: Diagnosis Step 2 One or more PES statements using NCPT terms Medical diagnoses; vague problem descriptions I: Intervention Step 3 Nutrition prescription, goals, delivery, education, counselling, coordination Data restated from assessment M/E: Monitoring and evaluation Step 4 Indicators to be measured, criteria for comparison, timeframe, follow-up plan Generic statements such as "will follow up" Source: Nutrition Care Process step definitions, Academy of Nutrition and Dietetics (ncpro.org); section content as taught in this guide. A good note is short and traceable. Every number in the assessment section should be there because it supports or excludes a diagnosis. Every diagnosis should trace back to numbers in the assessment. Every intervention should trace to an etiology. Every monitoring indicator should trace to a sign or symptom. If you can draw those arrows through your own note, it is a good note. If you cannot, no amount of additional data will rescue it. What the standardisation bought: a documented example The argument for standardised process and language is easy to state and hard to believe until you see it pay. The clearest documented example in American dietetics is the Malnutrition Quality Improvement Initiative, developed by the Academy of Nutrition and Dietetics with the consulting firm Avalere Health under funding from the Academy Foundation. Its premise was simple. Hospital malnutrition was known to be common and known to be under-recognised, but because identification and documentation varied so widely between clinicians and institutions, nobody could show whether better dietetic practice changed anything measurable. The initiative produced a toolkit and a set of electronic clinical quality measures covering screening within twenty-four hours of admission, completion of a nutrition assessment for screened-positive patients, documentation of a malnutrition diagnosis by the physician or advanced practice provider, and documentation of a nutrition care plan for patients identified as malnourished. Those measures are only constructible because the underlying process and terminology are standardised. You cannot write a quality measure on "the dietitian did a good job." You can write one on "a nutrition assessment using defined criteria was completed and documented within a defined window." The practical consequence for a student is worth absorbing. Malnutrition measures derived from this work entered national quality reporting in the United States, which means that the way you write a note is no longer a private matter between you and your preceptor. Aggregated across a department, it becomes evidence about whether nutrition care is happening at all. This is also why hospitals care intensely about whether the nutrition diagnosis is captured in a form the coding system can read, and why a beautifully reasoned care plan written in idiosyncratic language is, institutionally, close to invisible. Who the process serves A persistent misreading treats the Nutrition Care Process as a hospital tool. It is not. The Academy's model names individuals, groups and populations as recipients, and the terminology includes a population-based nutrition action intervention domain precisely because dietitians work in public health, community programmes, food service management, school systems and private practice as well as on wards. The process adapts rather than changing. In a community setting, the assessment leans more heavily on food and nutrition-related history and client history and much less on biochemical data, because there may be no laboratory values at all. The comparative standards shift from clinical reference ranges to dietary guidelines and population intake recommendations. The diagnosis is more likely to sit in the behavioural-environmental domain — food insecurity, limited access to food, undesirable food choices — than in the clinical domain. The intervention is more likely to be education, counselling or coordination of care than a change in the diet order. Monitoring may run over months rather than days. None of that is a different process. It is the same four steps with different parameters, which is exactly the point this guide keeps returning to: learn the structure once and the settings become variations rather than separate subjects. Where this leaves the textbook Krause and Mahan's clinical chapters are, viewed through this lens, an enormous reference for two of the four steps. They tell you what to assess in a given condition and what interventions the evidence supports. They tell you comparatively little about diagnosis, because diagnosis is a reasoning skill that cannot be tabulated by disease, and comparatively little about monitoring, because monitoring depends on what you diagnosed. This is why students who read the textbook well still write poor care plans. The two steps the book covers most thoroughly are the two that require the least judgement, and the two it covers most lightly are the ones examinations test hardest. The remainder of this guide is weighted accordingly. Key Takeaways · The Nutrition Care Process is a reasoning method adopted by the Academy of Nutrition and Dietetics in 2003; it standardises how dietitians think and record, not what any individual patient receives. · Its four steps are assessment, diagnosis, intervention, and monitoring and evaluation, and each step constrains the next in a chain that should be traceable in any note. · The Nutrition Care Process Model adds three things the steps alone do not show: the process is cyclical, it depends on the practitioner's own competencies, and it operates inside a health care and social environment that limits what is possible. · Screening identifies who enters the process; it is not one of the four steps. · The Nutrition Care Process Terminology, maintained in the eNCPT, supplies controlled vocabulary for each step and makes outcomes measurable across patients. · ADIME is the four steps written down; misplacing content between its sections usually reveals a reasoning error, not a formatting one. Review Questions 1. Explain why the Academy adopted a standardised process and terminology, and what specific problem in practice it was intended to solve. 1. Name the four steps of the Nutrition Care Process and state, for each, what constraint it places on the step that follows. 2. Why is nutrition screening not counted as a step of the Nutrition Care Process, and who typically performs it? 3. Describe three elements of the Nutrition Care Process Model that lie outside the four central steps, and give a clinical example of how each can change a care plan. 4. A student's ADIME note contains the sentence "Will recommend oral nutrition supplement twice daily" in the assessment section. What reasoning error does this most likely indicate? 5. Why did ADIME replace the SOAP note in most dietetic practice settings? Chapter 2: Assessment I — The Five Categories of Data Nutrition assessment is the step students find easiest to do badly while appearing to do well. Collecting data is comfortable work: there are forms, there are laboratory printouts, there are charts to copy from. A student can fill two pages of an assessment section and still have no idea what is wrong with the patient, because collecting is not assessing. Assessment is the comparison of collected data against a standard, and the interpretation of the gap. Krause and Mahan devotes substantial space to this step, and Raymond and Morrow structure it as the Academy does, around five categories of data. Learning those five categories as a checklist is the single most useful memorisation in clinical dietetics, because it prevents the most common assessment failure, which is not a wrong number but a missing one. Table 2 sets out the categories with their terminology codes. Table 2. The nutrition assessment data categories and their terminology codes. Code Category Representative content Typical source FH Food/nutrition-related history Intake, diet history, medication and supplement use, knowledge, beliefs, food availability, physical activity Patient interview, recall, food record AD Anthropometric measurements Height, weight, BMI, weight history, growth pattern, body composition Direct measurement, chart, growth chart BD Biochemical data, medical tests and procedures Laboratory values, gastric emptying studies, resting metabolic rate measurement Medical record, laboratory system PD Physical exam findings Muscle and fat wasting, oedema, oral health, skin, hair, nails, affect, vital signs Nutrition-focused physical examination CH Client history Medical and surgical history, social history, family history, treatment history Medical record, patient interview Source: Nutrition Care Process Terminology assessment domains, Academy of Nutrition and Dietetics (eNCPT, 2023 edition; ncpro.org). Two further categories appear in the current terminology and are often overlooked: comparative standards, which is what the data are measured against, and assessment tools, which covers validated instruments such as screening questionnaires and subjective global assessment. Comparative standards in particular deserves attention, because it is the category that turns collection into assessment. A serum albumin of 24 g/L is a datum. A serum albumin of 24 g/L against a reference range of 35 to 50 g/L, in a patient with C-reactive protein of 180 mg/L, is an assessment — and the assessment is about inflammation, not about nutrition. Food and nutrition-related history This is the category that is genuinely ours. Physicians order the laboratory tests and nurses record the weights, but nobody else in the building is going to find out what the patient actually eats. It is also the category most vulnerable to poor technique, because everything in it is reported rather than measured. The core methods are the twenty-four hour recall, the food frequency questionnaire, the food record or diary, and the diet history, and the textbook treats each with appropriate scepticism. The twenty-four hour recall is quick and imposes no burden on the patient, but a single day rarely represents habitual intake and the multiple-pass technique — prompting through the day several times, in increasing detail, with specific probes for forgotten items such as beverages, condiments and food eaten standing up — is what separates a usable recall from a useless one. Food frequency questionnaires capture habitual patterns over months but quantify poorly. Food records prospectively avoid memory error but change behaviour in the act of recording, usually downwards. Diet history interviews combine methods and take the longest. Underreporting is systematic rather than random, which is what makes it dangerous. It is more pronounced in people with higher body weight, in women, and in relation to foods the patient perceives as socially undesirable. The correct response is not to inflate reported intake by some guessed factor but to record what was reported, note the limitations of the method, and triangulate against weight trend, which is the only intake measure that never lies. If a patient reports 1,200 kcal daily and has gained four kilograms in three months, the report is wrong and the weight is right. This category also holds medication and supplement use, and this is where students most often fail to look. Food-drug interactions belong in the assessment: metformin and vitamin B12 absorption, proton pump inhibitors and B12, iron and magnesium, loop diuretics and potassium, magnesium and thiamin, long-term corticosteroids and bone, warfarin and vitamin K variability, and the growing relevance of GLP-1 receptor agonists, which reduce intake so effectively that they can produce protein and micronutrient inadequacy and loss of lean mass in patients whose weight loss looks like success. Herbal and dietary supplement use is underreported unless asked about directly and without judgement. Anthropometric measurements Height and weight sound trivial and are not. A stated height is often the height the patient was at thirty; vertebral compression and kyphosis in older adults can reduce measured stature by several centimetres, and since BMI divides weight by height squared, a height error propagates. Where standing height cannot be obtained, knee height or arm span equations are used, and the textbook covers them. Weight demands even more care in hospital. A patient in positive fluid balance can carry several kilograms of oedema or ascites; a weight recorded on a bed scale with linen, pumps and a limb immobiliser included is not comparable to yesterday's standing weight. The clinically meaningful figure is almost always the weight trend and the percentage change over a defined interval, not the absolute number. Percentage weight change is calculated as the difference between usual body weight and current weight, divided by usual body weight, multiplied by one hundred — and getting a reliable usual body weight from the patient is a skill, because "I've always been about eleven stone" needs converting, dating and cross-checking. Body mass index classifies but does not diagnose. This is not a minor caveat. In 2025 a Lancet Diabetes and Endocrinology Commission chaired by Francesco Rubino argued that BMI should be treated as a population-level risk marker rather than an individual diagnostic measure, and proposed distinguishing clinical obesity — excess adiposity with demonstrable organ or tissue dysfunction — from preclinical obesity, where adiposity is excessive but function is preserved. The Commission recommends confirming excess adiposity either by direct body fat measurement or by BMI combined with at least one anthropometric measure such as waist circumference, waist-to-hip ratio or waist-to-height ratio. For a student, the practical upshot is that a BMI alone is never an adequate anthropometric assessment. Body composition methods vary enormously in what they can deliver at the bedside. Bioelectrical impedance is cheap and portable but assumes normal hydration, which makes it unreliable in exactly the acutely ill patients whose muscle mass you most want to know. Dual-energy X-ray absorptiometry is accurate and rarely available for nutrition purposes. Computed tomography images taken for other reasons can be analysed at the third lumbar vertebra to quantify skeletal muscle, and this is increasingly used in oncology research. Mid-upper arm circumference and calf circumference are crude but need nothing but a tape measure, and calf circumference below 31 centimetres is a widely used indicator of reduced muscle mass in older adults. Biochemical data, medical tests and procedures The laboratory section of an assessment is where students most often mistake volume for insight. The discipline to acquire is asking, of every value you copy across, what nutritional question it answers and what else could explain it. The most important correction to make early concerns the visceral proteins. Albumin and prealbumin — transthyretin — were taught for decades as nutritional markers and are still described that way in older resources. They are negative acute-phase reactants. In inflammation, hepatic synthesis shifts, capillary permeability rises, and concentrations fall regardless of nutritional intake; they also fall with fluid overload, liver disease and nephrotic protein loss, and rise with dehydration. The Academy and the American Society for Parenteral and Enteral Nutrition both explicitly excluded them from the adult malnutrition characteristics for this reason. They remain useful as markers of inflammatory burden and of illness severity, and a falling albumin in a patient who is clinically stable should prompt a search for a new inflammatory process rather than an increase in the protein prescription. What is genuinely informative varies with the question. Haemoglobin, mean corpuscular volume, ferritin interpreted alongside C-reactive protein, transferrin saturation, serum B12 with methylmalonic acid where B12 is borderline, and red cell folate answer questions about specific deficiencies. Urea and creatinine, electrolytes, and estimated glomerular filtration rate define the renal constraints on a prescription. Glucose and glycated haemoglobin define the glycaemic ones, with the caveat that glycated haemoglobin is unreliable in haemolysis, recent transfusion, iron deficiency and advanced kidney disease. Liver enzymes and bilirubin, alongside international normalised ratio and ammonia, define hepatic ones. Phosphate, potassium and magnesium are the refeeding triad and must be checked before nutrition support is started in anyone who has been underfed. Twenty-five-hydroxyvitamin D, calcium corrected for albumin, and parathyroid hormone answer bone questions. Zinc and copper are worth measuring in malabsorption, burns, high-output fistulae and long-term parenteral nutrition, with the caution that zinc also falls in inflammation. Nutrition-relevant procedures belong here too: swallow studies, gastric emptying scintigraphy, endoscopy and biopsy findings, and indirect calorimetry, which is the measured alternative to the predictive equations discussed in the next chapter. Nutrition-focused physical findings The nutrition-focused physical examination is the part of assessment that most reliably separates a confident practitioner from a hesitant one, and it is the part students get least practice in. It is a structured, hands-on examination of the areas where nutritional depletion shows: subcutaneous fat at the orbital region, the triceps and the ribs; muscle at the temples, the clavicles, the shoulders, the interosseous region of the hand, the quadriceps and the calves; fluid accumulation at the ankles, sacrum and abdomen; and the micronutrient-sensitive tissues — skin, hair, nails, eyes, tongue, gums and lips. The findings map onto deficiencies in ways worth knowing. Angular cheilitis and glossitis suggest B vitamin inadequacy. Spoon-shaped nails suggest iron deficiency. Perifollicular haemorrhage and corkscrew hairs suggest vitamin C deficiency. Bitot's spots and night blindness suggest vitamin A deficiency. Dermatitis in a photosensitive distribution alongside diarrhoea and cognitive change suggests niacin deficiency. A flapping tremor points to hepatic encephalopathy rather than a nutrient. The examination also includes functional measures, most usefully handgrip strength by dynamometer, which changes earlier than body composition does and is one of the six characteristics in the Academy and American Society for Parenteral and Enteral Nutrition malnutrition criteria. Two disciplines make the examination useful. The first is always examining bilaterally and comparing with what you would expect for that person's age, sex and usual build rather than against an abstract norm. The second is grading findings as mild, moderate or severe rather than simply present or absent, because severity determines the severity of the malnutrition diagnosis that may follow. Client history Client history is the category students treat as background and clinicians treat as decisive. It holds medical and surgical history, which tells you what physiology is altered; treatment history, including radiation fields, chemotherapy agents and their known nutritional toxicities, and previous nutrition support; family history; and social history — living situation, who cooks, cooking facilities, income, food assistance programme participation, education level, language, cultural and religious food practices, substance use, and social support. The social history is not sentiment. It is what determines whether the intervention will happen. A patient with stage four chronic kidney disease living alone with no cooking facilities and no refrigerator cannot execute a low-potassium, low-phosphorus cooking plan, and a dietitian who writes one has produced a document rather than a plan. Raymond and Morrow return to social determinants throughout the clinical chapters, and the 2024 revision of the Academy's Scope and Standards of Practice makes striving for health equity, including attention to food security and social determinants, an explicit standard rather than an optional consideration. Screening, and why it sits outside the process Screening identifies who needs a full assessment, and in most institutions it must be completed within twenty-four hours of admission. The tools are deliberately brief and are usually completed by nursing staff: the Malnutrition Screening Tool asks about unintentional weight loss and reduced appetite in two questions; the Malnutrition Universal Screening Tool combines BMI, unplanned weight loss and the effect of acute disease on intake; Nutritional Risk Screening 2002 adds severity of illness and age. Each has a defined cut-off above which referral for assessment follows. The essential distinction is that screening is designed to be sensitive rather than specific. It over-identifies deliberately, because the cost of a false positive is a dietitian's time and the cost of a false negative is a missed malnutrition diagnosis. A positive screen is therefore not a diagnosis and never appears in the diagnosis section of a note; it appears in the assessment section as the reason the patient was referred. A workable sequence Students are rarely taught an order in which to do this, and the order matters because time on a ward is short and the interview is the part most easily truncated. Work the record before the patient. Read the admission note, the problem list, the operative record if there is one, the medication chart and the laboratory trends; extract the client history, biochemical data and whatever anthropometrics are recorded. This takes ten minutes and means you arrive at the bedside already knowing what you are looking for. It also means you can ask precise questions — "the notes say you had part of your small bowel removed in March, how have your bowels been since?" — rather than generic ones, which patients answer generically. Then take the food and nutrition-related history at the bedside, because it needs the patient's attention and cannot be reconstructed afterwards. Then perform the physical examination, which is best done at the end of the encounter when rapport exists and touching the patient's shoulders and hands is no longer an intrusion. Measure or verify height and weight yourself where the patient's condition allows. Only then do the arithmetic: estimated energy, protein and fluid requirements, percentage weight change, BMI, and comparison of reported intake against estimated needs as a percentage. That arithmetic is the subject of the next chapter, and it is the point at which the assessment finally answers a question rather than merely describing a person. One habit worth building immediately. As you work, keep a running list of candidate diagnoses and the evidence for and against each. By the time the arithmetic is done, you will usually find that two or three candidates survive and one dominates, and the diagnosis step becomes a matter of choosing rather than of inventing. Key Takeaways · Assessment is comparison against a standard, not collection; the comparative standards category is what converts data into interpretation. · The five categories — food/nutrition-related history, anthropometrics, biochemical data, physical findings and client history — work best as a checklist against omission. · Underreporting of intake is systematic, not random; weight trend is the check on every reported intake. · Albumin and prealbumin are negative acute-phase reactants, not nutritional markers, and were deliberately excluded from the consensus malnutrition characteristics. · The nutrition-focused physical examination and handgrip strength provide findings no laboratory value can supply, and must be graded for severity. · Screening is sensitive by design, sits outside the four steps, and belongs in the assessment section of a note rather than the diagnosis section. Review Questions 1. Distinguish collecting data from performing an assessment, and explain the role of comparative standards in that distinction. 1. A patient reports an intake of 1,100 kcal per day but has gained 3 kg over two months. What does this tell you, and how should you document it? 2. Explain why serum albumin fell out of use as a nutritional marker, and state two situations in which it is still clinically informative. 3. List the body regions examined in a nutrition-focused physical examination for fat loss, muscle loss and fluid accumulation, and explain why findings are graded rather than simply recorded as present. 4. Why are screening tools designed to be more sensitive than specific, and what error does a student make by recording a positive screen as a nutrition diagnosis? 5. Give three items from the client history category that could make an otherwise appropriate nutrition prescription undeliverable. Chapter 3: Assessment II — Estimating Needs and Identifying Malnutrition An assessment that does not end in numbers has not ended. Estimating energy, protein and fluid requirements is the arithmetic that converts a description of a patient into a target you can prescribe against and measure towards, and identifying malnutrition is the judgement that most often determines whether a nutrition diagnosis is made at all. Both are places where students lose marks for avoidable reasons: using an equation outside the population it was derived in, applying a stress factor without saying why, or conflating two different consensus frameworks for malnutrition into a single half-remembered list. Energy: measurement first, prediction second The gold standard for determining energy requirements is measurement, and the measurement technique is indirect calorimetry. It works from a simple relationship: the oxidation of carbohydrate, fat and protein consumes oxygen and produces carbon dioxide in ratios characteristic of each substrate, so measuring inspired and expired gas volumes and concentrations over a period allows resting energy expenditure to be calculated, conventionally through the abbreviated Weir equation. The respiratory quotient — carbon dioxide produced divided by oxygen consumed — is reported alongside, and in a validly performed test lies roughly between 0.67 and 1.3. Values outside that range indicate a technical problem, not an exotic metabolism, and the commonest causes are air leaks around a mask or tracheal tube cuff, a patient who has not rested, and high inspired oxygen concentrations. The test has practical requirements that examinations like to ask about: the patient should be rested, ideally for at least thirty minutes in a thermoneutral quiet environment, fasted or at least in a steady feeding state, without recent physiotherapy, and clinically stable with inspired oxygen below about sixty per cent and no significant air leak. Where measurement is unavailable — which is most of the time — prediction is used, and the equation most widely recommended for healthy non-obese and obese adults is Mifflin-St Jeor. It estimates resting metabolic rate as ten times weight in kilograms, plus 6.25 times height in centimetres, minus five times age in years, plus five for men and minus 161 for women. Its standing comes from a systematic review by Frankenfield and colleagues published in the Journal of the American Dietetic Association in 2005, which compared the available equations and found Mifflin-St Jeor the most reliable in both non-obese and obese healthy adults; it remains the Academy's recommended default in that population. Several cautions travel with it. The equation predicts resting metabolic rate, not total energy expenditure, so an activity factor must be applied — conventionally around 1.2 for a bedbound or minimally mobile patient rising to 1.6 or above with substantial activity. It was derived in healthy adults, not in acute illness, and its accuracy in critical illness is poor. It is one prediction from a distribution: even in the population where it performs best, a substantial minority of individuals fall more than ten per cent outside the predicted value, which is why the estimate must be revised against observed weight response rather than defended. In acute and critical illness the field has moved towards simple weight-based estimates rather than equations plus stress factors. The 2022 American Society for Parenteral and Enteral Nutrition guideline for the critically ill adult, led by Charlene Compher, recommends feeding between 12 and 25 kcal per kilogram in the first seven to ten days of intensive care, noting that trials comparing higher and lower energy delivery in that window have not demonstrated a difference in outcome. That recommendation carries a message students should absorb: in the early phase of critical illness, precision about the energy target matters less than avoiding both starvation and overfeeding, and endogenous substrate mobilisation means that full energy provision early may not be beneficial. In obesity, hypocaloric high-protein feeding is a recognised approach in critical care, and where predictive equations are used in obesity it is conventional to use actual body weight in Mifflin-St Jeor rather than an adjusted weight, because the equation was validated in obese as well as non-obese subjects. Practices differ between institutions and the textbook covers the alternatives; what matters for an examination is that you state which weight you used and why. Protein and fluid Protein requirements are expressed per kilogram and the reference point is the Recommended Dietary Allowance of 0.8 grams per kilogram per day for healthy adults, which represents the minimum sufficient to meet the needs of nearly all healthy people rather than an optimum. Clinical requirements rise above it in almost every acute condition, and the textbook's condition-specific chapters supply the figures: roughly 1.2 to 2.0 grams per kilogram per day in critical illness per the American Society for Parenteral and Enteral Nutrition 2022 guideline, 1.0 to 1.2 grams per kilogram in maintenance dialysis per the 2020 Kidney Disease Outcomes Quality Initiative update, and restricted intakes of 0.55 to 0.60 grams per kilogram in non-dialysis chronic kidney disease without diabetes under that same guideline. There is also a substantial literature arguing that older adults benefit from intakes above the Recommended Dietary Allowance, in the region of 1.0 to 1.2 grams per kilogram, to counter anabolic resistance. Which weight to use is a recurring student error. The general convention is actual body weight in patients within a normal weight range, ideal or adjusted body weight in substantial obesity, and dry or oedema-free weight in fluid overload — and in every case the note should say which was used. An enormous protein target calculated on a weight inflated by twelve litres of ascites is not a rigorous estimate; it is a mistake dressed up in decimal places. Fluid requirements are commonly estimated at around 30 to 35 millilitres per kilogram per day in adults, or by the Holliday-Segar method based on weight bands, which allows 100 millilitres per kilogram for the first ten kilograms, 50 millilitres per kilogram for the next ten, and 20 millilitres per kilogram thereafter. Both are starting points to be adjusted for losses — fever, high-output stomas, fistulae, drains, burns, diarrhoea — and for restriction in heart failure, advanced kidney disease, syndrome of inappropriate antidiuretic hormone secretion and ascites. In tube-fed patients the water content of the formula and of flushes must be counted, and in parenteral nutrition the volume of the admixture is itself part of the fluid balance. Anthropometric interpretation across the lifespan Adult anthropometric interpretation revolves around BMI categories, weight change, and increasingly waist circumference, for the reasons set out in the previous chapter. Paediatric interpretation revolves around growth charts, and using the wrong chart is an error that changes the answer. The Centers for Disease Control and Prevention recommends the World Health Organization Child Growth Standards for infants and children from birth to two years, and the CDC 2000 growth charts for children and adolescents from two years onwards. The distinction matters because the two are different kinds of instrument. The WHO charts are standards, derived from a multicentre study of children raised under conditions considered optimal, including breastfeeding, and describe how children should grow. The CDC charts are references, derived from survey data on how American children did grow in a period of rising obesity prevalence, and describe a population rather than an ideal. In 2022 the CDC published extended BMI-for-age growth charts adding percentile curves above the ninety-fifth — the ninety-eighth, ninety-ninth, 99.9th and 99.99th — and extending plottable BMI up to 60 kilograms per square metre. They are recommended for children aged two and over whose BMI lies above the ninety-seventh percentile, and they exist because the original charts could not usefully track children with very high BMI, whose values simply ran off the top of the page. Reading a growth chart well means reading the trajectory, not the point. A child steadily tracking the tenth percentile is usually a small child. A child who has fallen from the fiftieth to the tenth across two centile spaces in six months is a clinical problem, whatever the current percentile says. Crossing centile lines downwards in weight-for-age, then weight-for-length or BMI-for-age, then eventually length or height-for-age, is the classic sequence of faltering growth, and the involvement of linear growth indicates chronicity. Identifying malnutrition: two frameworks, kept separate Malnutrition identification is where clinical assessment becomes consequential, because it determines whether a diagnosis is made, how the case is coded, and in many systems whether the admission is reimbursed at a higher acuity. Two frameworks dominate, and students routinely merge them into a single confused list. Keep them apart. The first is the 2012 consensus statement of the Academy of Nutrition and Dietetics and the American Society for Parenteral and Enteral Nutrition, led by Jane White, which set out six characteristics for identifying and documenting adult malnutrition: insufficient energy intake, weight loss, loss of body fat, loss of muscle mass, localised or generalised fluid accumulation that may mask weight loss, and diminished functional status as measured by handgrip strength. Two or more of the six are required, and the diagnosis is made within one of three contexts — acute illness or injury, chronic illness, or social or environmental circumstances including starvation — each with its own thresholds for non-severe and severe malnutrition. These are commonly abbreviated as the AAIM criteria. The second is the Global Leadership Initiative on Malnutrition framework, published in 2019 by Cederholm, Jensen and colleagues on behalf of the major global clinical nutrition societies, and updated in a five-year review led by Jensen in 2025. GLIM uses a two-step model: validated screening first, then diagnostic assessment requiring at least one phenotypic and at least one etiologic criterion. The phenotypic criteria are weight loss, low body mass index, and reduced muscle mass; the etiologic criteria are reduced food intake or assimilation, and disease burden with inflammation. Severity is graded by the phenotypic criteria alone into stage 1, moderate, and stage 2, severe. Table 3 sets the two side by side. Table 3. Two consensus frameworks for identifying adult malnutrition. Feature AAIM (Academy/ASPEN, 2012) GLIM (2019; updated 2025) Structure Two or more of six characteristics, within a stated context One phenotypic plus one etiologic criterion, after positive screening Weight-loss threshold, severe 5% in 1 month or 7.5% in 3 months, context dependent More than 10% within 6 months, or more than 20% beyond 6 months Low BMI Not a listed characteristic Under 20 if aged under 70; under 22 if 70 or over Muscle and fat Loss of muscle mass and loss of body fat, graded by physical exam Reduced muscle mass by validated body composition technique Inflammation Expressed through the three diagnostic contexts An explicit etiologic criterion Functional measure Reduced handgrip strength is one of the six Supportive, not a diagnostic criterion Sources: White JV et al., J Acad Nutr Diet 2012;112(5):730-738; Cederholm T et al., Clin Nutr 2019;38(1):1-9; Jensen GL et al., Clin Nutr and JPEN 2025 five-year update. The frameworks disagree in ways that matter. A 2025 Perspectives in Practice article in the Journal of the Academy of Nutrition and Dietetics by Compher, Jensen, Aloupis and Steiber made the point sharply: GLIM does not classify weight loss below ten per cent as severe, whereas the Academy and American Society for Parenteral and Enteral Nutrition criteria identify losses of five to 7.5 per cent as severe depending on the interval and context. A patient can therefore be severely malnourished by one framework and moderately malnourished by the other on identical data, with real consequences for coding and for insurance claims. Rather than endorsing one, those authors describe the frameworks as complementary and recommend documenting the variables required by both during assessment. That is sound advice for a student too: collect weight history, BMI, intake history, muscle and fat assessment, fluid status, functional measures and inflammatory context, and you can populate either framework. Two further points. First, both frameworks are identification tools, not nutrition diagnoses in the terminology sense: the corresponding standardised diagnosis is malnutrition in the clinical domain, written as a PES statement in the usual way. Second, the 2025 GLIM update reaffirmed the core criteria while strengthening guidance on muscle mass measurement and on assessing inflammation, and validation work continues; the framework should be expected to evolve, and the edition of any guideline you cite should be stated. Sarcopenia, cachexia and the distinctions that get blurred Malnutrition is one of several overlapping states of nutritional and muscular depletion, and being able to distinguish them is worth marks and, more importantly, changes what the intervention can realistically achieve. Sarcopenia is the loss of skeletal muscle mass and function. The revised European consensus published by Cruz-Jentoft and colleagues in 2019 reoriented the definition around muscle strength as the primary indicator, with low muscle quantity or quality confirming the diagnosis and low physical performance indicating severity. It is age-associated but not exclusively so, and the intervention that addresses it is resistance exercise combined with adequate protein and energy, not protein alone. A nutrition plan for a sarcopenic older adult that does not mention activity has missed half the treatment. Cachexia is different again: a multifactorial syndrome of ongoing loss of skeletal muscle, with or without fat loss, driven by an underlying illness and its inflammatory response, and not fully reversible by conventional nutrition support. Cancer cachexia is the most studied form, with an international consensus definition published by Fearon and colleagues in 2011 staging it as precachexia, cachexia and refractory cachexia. Heart failure, chronic obstructive pulmonary disease, chronic kidney disease and rheumatoid arthritis all have recognised cachectic forms. The clinically important consequence is that the goal of nutrition intervention changes with the state. Starvation-related malnutrition without inflammation can be substantially reversed by feeding. Malnutrition on a background of chronic inflammation responds partially. Refractory cachexia does not respond, and pushing energy intake in that situation adds burden without benefit, which is why nutrition care in advanced disease shifts towards symptom relief, food enjoyment and the relief of family distress about eating. Recognising which state you are in is therefore an assessment task with direct consequences for what you may honestly promise. Key Takeaways · Indirect calorimetry measures resting energy expenditure; predictive equations estimate it, and Mifflin-St Jeor is the recommended default in healthy non-obese and obese adults on the strength of Frankenfield and colleagues' 2005 systematic review. · Predictive equations give resting metabolic rate, require an activity factor, perform poorly in acute illness, and must be revised against observed weight response. · Protein and fluid estimates must state which body weight was used; a target calculated on an oedematous weight is a mistake, not a rigorous estimate. · WHO growth standards are used from birth to two years and CDC references from two years onwards; the 2022 CDC extended BMI charts exist for children above the ninety-seventh percentile. · Growth charts are read as trajectories; crossing centiles downwards, eventually involving linear growth, is the pattern of faltering growth. · AAIM and GLIM are distinct frameworks with different structures and different severity thresholds; collect the variables for both and state which framework a conclusion rests on. Review Questions 1. State the Mifflin-St Jeor equation for men and for women, explain what it predicts, and name two adjustments needed before it becomes a usable energy target. 1. A respiratory quotient of 1.45 is reported from an indirect calorimetry study. What should you conclude and what would you check? 2. Which body weight would you use to estimate protein needs in a patient with a BMI of 41 and four litres of ascites, and how would you document the choice? 3. Explain the difference between a growth standard and a growth reference, and state which chart applies at 18 months and which at 6 years. 4. Compare the AAIM and GLIM frameworks on structure, the role of inflammation, and the threshold for severe weight loss. Describe a patient who would be graded differently by each. 5. Why is the identification of malnutrition by either framework not, by itself, a nutrition diagnosis? Hashtags: #TheDieteticBlueprint #KrauseAndMahan #NutritionCareProcess #ClinicalDietetics #ADIME #NutritionAssessment #NutritionDiagnosis #NutritionIntervention #MonitoringAndEvaluation #NutritionCareProcessTerminology #eNCPT #PESStatement #ComparativeStandards #NutritionFocusedPhysicalExam #MedicalNutritionTherapy #MalnutritionScreening #MalnutritionDiagnosis #AAIMCriteria #GLIMCriteria #IndirectCalorimetry #MifflinStJeorEquation #Sarcopenia #Cachexia #EvidenceBasedDietetics #FutureOfClinicalNutrition
- The Digital Archive (Memory, Erasure, and Curatorial Power in Cyberspace)
Download the Book (PDF): Introduction In the autumn of 2009, Yahoo shut down GeoCities. For fifteen years the service had given away small plots of web space to anyone who wanted them, and millions of people had taken up the offer. They built pages about their cats, their churches, their illnesses, their favourite bands, their grief. They wrote poetry in blinking text and pasted in animated construction signs to announce that the page was still being built. None of it was important in the way that treaties and parliamentary debates are important. All of it was evidence of how ordinary people first learned to speak in public through a new medium. When Yahoo announced the closure, the company gave its users a few months' notice. A loose collective of volunteers calling themselves Archive Team, led by the historian of computing Jason Scott, began downloading as much of GeoCities as they could reach before the servers went dark. They saved a great deal. They did not save it all. What they saved exists because a handful of people decided, on their own time and without anyone's permission, that it mattered. That episode contains, in miniature, the argument of this book. The first thing it shows is that digital material is not durable by default. It survives only when someone pays for the servers, maintains the formats, renews the domain names, and decides not to switch it off. The second is that the decision about what survives is rarely made by the people whose lives the material records. It is made by corporations weighing costs, by states weighing risks, by institutions weighing mandates, and increasingly by software weighing signals of relevance. The third is that the rescue itself was a curatorial act. Archive Team's crawlers followed links, and what was poorly linked was more likely to be lost. The archive of GeoCities that exists today is a particular shape, cut by the tools that made it and the urgency under which they ran. The argument This book argues a single proposition. The digital archive is not a storehouse in which the past waits to be found. It is a continuous exercise of power over what the past is permitted to become. Every stage in the life of a digitized record, from the choice to scan one collection rather than another, through the metadata that describes it, the servers that hold it, the licences that fence it, the search engines that rank it and the policies that delete it, is a decision with winners and losers. Those decisions are mostly invisible, and their invisibility is the source of much of their force. The practical consequence is that stewardship of digital memory has to be judged not by how much it stores but by whether the choices it makes can be seen, contested and corrected by the people they affect. That claim cuts against two popular intuitions. The first is that the internet remembers everything, a belief captured in the familiar warning that whatever you post will follow you forever. There is truth in it for certain kinds of content, especially embarrassing personal material copied across many sites. But as a general description of the web it is wrong. A study published by the Pew Research Center in 2024 found that a substantial share of web pages that existed in 2013 could no longer be reached a decade later, and that link rot had eaten into news articles, government pages and the references cited on Wikipedia. The web forgets constantly. What it remembers is not random, and the pattern of its remembering is one of this book's subjects. The second intuition is that digitization is a form of democratization: that once the manuscripts, newspapers, photographs and recordings of the past are scanned and posted online, they belong to everyone. Here too there is something real. A student in Lagos can now read colonial-era newspapers that once required a flight to London and a reader's ticket at the British Library. But she may find that those newspapers sit behind a commercial paywall her university cannot afford, that the search tool that indexes them was trained on English typefaces and misreads the local-language columns, and that the collection she can reach was selected, decades ago, by officials whose purpose was to govern her great-grandparents. Access has widened. The terms of access have not become neutral. Why the archive is a political object Archives have always been political. Ancient rulers kept tax registers and king lists; early modern states kept censuses and passport records; colonial administrations kept files on the populations they ruled and, in many cases, destroyed or removed those files when they left. Scholars across several disciplines have spent the last half-century making this visible. The Haitian anthropologist Michel-Rolph Trouillot showed how silences enter history at every stage of its making. Jacques Derrida traced the word archive back to the arkheion, the house of the magistrates who held both the documents and the authority to interpret them. Ann Laura Stoler taught historians to read colonial files not only for what they said but for the anxieties that shaped what officials thought worth writing down. The digital turn did not end these dynamics. It multiplied and disguised them. When a physical archive excludes a document, the exclusion often leaves traces: a gap in a numbered series, a note in a register, a destruction certificate. When a search algorithm ranks a document on the fortieth page of results, nothing records the fact that almost nobody will ever see it. When a platform removes a video under a content policy, the video can vanish from the record entirely, along with the evidence that it ever existed. When a server migration corrupts a decade of uploads, as happened to MySpace in 2019 when the company acknowledged it had lost music uploaded over roughly twelve years, the loss is announced, if at all, in a brief statement. The mechanisms of forgetting have become more efficient and less legible at the same time. There is also a new category of actor. Traditional archives were run by states, churches, universities and families. The digital archive is additionally run by technology companies whose primary obligation is to shareholders, by volunteer collectives with no formal mandate, by nonprofit libraries operating in legal grey zones, and by machine-learning systems that ingest the archived web as training material and return it to users in digested form. Each of these actors exercises curatorial power, and each is accountable, if at all, through a different mechanism: markets, courts, donors, or nothing. The shape of the book The chapters that follow move from foundations to mechanisms to ethics. Chapter 1 sets out the theoretical inheritance, showing why the archive was never neutral and how the colonial case, in particular the story of Britain's so-called migrated archives, makes the abstractions concrete. Chapter 2 examines what digitization actually changes, through the history of mass scanning projects from Google Books to Europeana and the less celebrated labour of metadata and optical character recognition. Chapter 3 turns to fragility: the bit rot, format obsolescence and link rot that threaten digital records, and the institutional arrangements built to resist them. Chapter 4 considers who keeps the web, focusing on the Internet Archive and the legal battles that have tested whether a nonprofit library can do for digital culture what public libraries did for print. Chapter 5 examines algorithmic curation, arguing that ranking is a form of remembering and that the systems which now mediate access to the past, including the large language models trained on it, embed particular judgements about relevance and authority. Chapter 6 turns to erasure by design: the right to be forgotten, platform content moderation, and the removal of government information, each of which deletes material for reasons that may be good or bad but that are rarely open to public scrutiny. Chapter 7 examines institutional gatekeeping, from commercial vendors who digitize public-domain material and sell it back to universities, to Indigenous communities who have built systems to control access to their own cultural heritage on their own terms. Chapter 8 addresses the most difficult ethical question in the field: what to do with records that are contested, painful or dangerous, from the files of secret police to photographs of lynchings to the daguerreotypes of enslaved people that were the subject of a long lawsuit against Harvard University. Chapter 9 draws the argument together into a set of principles for an accountable archive. The conclusion asks what follows for the rest of us, the people whose lives are being recorded, sorted and occasionally deleted by systems we did not design. A note on scope The subject is vast, and depth on a narrow front serves it better than a survey. The cases here are drawn mainly from North America, Europe and the former British Empire, with significant attention to Africa, the Middle East and Indigenous Australia and North America. That emphasis reflects where the best-documented controversies have occurred and where the literature is richest in English; it is not a claim that these regions matter most. The book treats archives in a broad sense, including libraries, museums, platforms and datasets, because the digital environment has blurred the boundaries between them. Archivists will notice that the term is used more loosely here than they would use it among themselves. The looseness is deliberate. The public now encounters its past through all these institutions at once, and the questions of power this book raises apply to all of them. One final orientation. It would be easy to read the argument as a counsel of despair, a catalogue of losses and abuses. That is not the intention. Many of the people who appear in these pages, from volunteer web archivists to Guatemalan records workers to Aboriginal community curators, have shown that digital memory can be held in trust rather than merely held. The point of exposing curatorial power is not to abolish it, which is impossible, but to insist that it be exercised in the open, by people who can be asked to account for their choices. Every archive decides. The question is who decides, how, and whether anyone else gets a say. Chapter 1: The Archive Was Never Neutral In 2011 five elderly Kenyans brought a claim in the High Court in London. They said that British colonial forces had tortured them during the Emergency of the 1950s, when the colonial government fought the insurgency it called Mau Mau. The British government's first line of defence was that liability had passed to the Kenyan state at independence. Its second, implicitly, was the state of the documentary record: the claimants had their bodies and their memories, but where were the files? Then, during the litigation, the Foreign and Commonwealth Office acknowledged that it held a large cache of colonial records at Hanslope Park, a secure government site in Buckinghamshire. The historian David Anderson, acting as an expert witness, had pressed for their disclosure. The cache turned out to contain material removed from Kenya and dozens of other territories on the eve of independence, kept out of the public archive for half a century. Among the papers were records of the procedures, known in some colonies as Operation Legacy, by which officials had sorted documents before handing power over: some to be left for the incoming government, some to be shipped to Britain, and some to be burned or dumped at sea. The case settled in 2013, with the British government expressing regret and paying compensation. The documents were transferred in stages to the National Archives at Kew, where they now form the series catalogued as FCO 141. Historians such as Caroline Elkins, whose book Imperial Reckoning had been attacked for relying heavily on oral testimony, found that the newly released files confirmed much of what the survivors had said. It is hard to imagine a clearer demonstration that archives are not neutral repositories. The record of colonial Kenya had been shaped, deliberately and at the highest levels, to protect the reputation of the power that created it. The official record was incomplete in a way that helped the state that made it, and that fact stayed hidden for decades. This chapter sets out the ideas that make such episodes intelligible, and it argues that those ideas matter even more once archives become digital. The theorists who first showed that archives are sites of power were writing about paper. Their insights have to be carried forward, because digital technologies have changed how archival power works without reducing how much of it there is. The custodian and the appraiser For much of the twentieth century the dominant professional self-image of the archivist was that of an impartial custodian. The English archivist Hilary Jenkinson, whose Manual of Archive Administration first appeared in 1922, argued that archives were the natural residue of administrative activity. The archivist's duty was to guard that residue, preserve its original order and keep his own judgement out of it. Records were evidence precisely because nobody had selected them for posterity; they had accumulated as a by-product of doing business. To interfere would be to contaminate the evidence. This was never a complete account even of paper archives, and it came under strain as the volume of modern government records exploded. By the middle of the century no institution could keep everything. The American archivist Theodore Schellenberg, writing in the 1950s, set out a theory of appraisal in which archivists deliberately chose what to keep according to its evidential and informational value. Appraisal made the archivist's judgement unavoidable and explicit. Someone had to decide which case files from a welfare office or a police department would survive, and the decision would shape what future historians could know. From the 1980s onward, archival theorists such as Terry Cook in Canada pushed further. Cook argued that archivists were active co-creators of the record rather than passive recipients of it. With Joan Schwartz he wrote an influential essay in the journal Archival Science in 2002 arguing that archives are sites where power is exercised: over what is recorded, what is kept, how it is described and who can see it. Randall Jimerson's Archives Power, published by the Society of American Archivists in 2009, brought the argument to a professional audience and asked archivists to accept responsibility for the social consequences of their choices. By the early twenty-first century few serious archivists still believed in the pure custodian. The profession had accepted, in principle, that it held power. The digital archive has in some respects reopened the custodial fantasy. When storage is cheap, it can seem that appraisal is unnecessary: why choose, when you can keep everything? The rhetoric of "big data" and of comprehensive capture, from web crawls that aim to preserve the whole public web to social media archives that aimed to hold every public post, revived the idea that the archive might simply record what happened without anyone choosing. As later chapters show, this is an illusion. Comprehensive capture is never comprehensive, and the choices that shape it have migrated from the archivist's desk into the design of crawlers, databases and ranking systems, where they are harder to see. The philosophers' archive Outside the archival profession, a separate line of thought treated the archive as a concept rather than an institution. In The Archaeology of Knowledge, published in French in 1969, Michel Foucault used the word archive to mean something close to the system of rules that governs what can be said in a given period: not a building full of documents, but the conditions that make certain statements possible and others unthinkable. This was a metaphorical extension, and archivists have sometimes complained that it confused debate. But it captured something real. What gets recorded depends on what the recording institutions consider sayable, and those limits are historically specific. Jacques Derrida's lecture Archive Fever, delivered in 1994 and published in English in 1996, began with etymology. The Greek arkheion was the residence of the archons, the magistrates, who held official documents and also had the right to interpret them. From the start, Derrida suggested, the archive joined two powers: the power to keep and the power to say what the kept things meant. He also emphasised that archiving is always oriented toward the future. We record in order to be remembered in particular ways, and the technology of recording shapes what can be recorded. Derrida, writing just as email was becoming widespread, speculated that psychoanalysis itself would have developed differently had Freud and his contemporaries corresponded by electronic mail. The medium is not a transparent container. These arguments can seem remote from practical questions about servers and metadata. They are not. When a technology company decides that its search engine will rank results by a measure of popularity, it is exercising the archontic power Derrida described: holding the record and determining how it will be read. When a platform's terms of service define what may be uploaded, it is setting the limits of what can be said in Foucault's sense, at least within its domain. The philosophers' archive turns out to be a good description of how digital platforms work. Silences at every stage The most useful single framework for thinking about archival power was offered not by a philosopher or an archivist but by an anthropologist and historian. In Silencing the Past, published in 1995, Michel-Rolph Trouillot argued that silences enter the production of history at four distinct moments. The first is the moment of fact creation, when sources are made: some events are recorded and others are not, and some people are in a position to write while others are not. The second is fact assembly, when sources are gathered into archives. The third is fact retrieval, when historians and others go to the archive and construct narratives from what they find. The fourth is the moment of retrospective significance, when a society decides which stories matter and how they fit into its larger account of itself. Trouillot's central example was the Haitian Revolution, the only successful slave revolt to found an independent state. He showed that European and North American contemporaries found the revolution literally unthinkable, because it contradicted their beliefs about race, and that the silence persisted in later scholarship, which treated the revolution as a footnote to the French one. The silence was not produced by a single act of suppression. It was cumulative, built up at each of the four moments by people who mostly did not intend to silence anything. The value of the framework for this book is that each of Trouillot's moments has a digital equivalent, and each is exposed to new forms of power. As Table 1 sets out, the stages of silencing map onto the lifecycle of digital records in ways that help locate where curatorial decisions are being made. Table 1. Trouillot's four moments of silence and their digital counterparts. Trouillot's moment Digital counterpart Where power sits Fact creation Born-digital recording and platform design Platform owners, device makers Fact assembly Digitization selection, web crawling, data retention Funders, vendors, crawler design Fact retrieval Search, ranking, recommendation, AI summaries Algorithm designers, interface owners Retrospective significance Virality, citation, training-data inclusion Networks of attention, model builders At the moment of fact creation, digital platforms determine what kinds of record can exist at all. A service that stores text but not images, or that deletes messages after twenty-four hours, shapes the record before any archivist is involved. At the moment of assembly, decisions about which collections to digitize and which websites to crawl act as a filter every bit as powerful as colonial officials sorting files for burning, though usually less malicious. At the moment of retrieval, search engines and discovery systems decide which documents a researcher will see first, and in practice which she will see at all. At the moment of retrospective significance, the dynamics of online attention and, increasingly, the composition of datasets used to train artificial intelligence determine which parts of the past are amplified and which fade. Reading along and against the grain If archives are shaped by power, how should anyone read them? The anthropologist Ann Laura Stoler, in Along the Archival Grain, published in 2009, offered one answer drawn from years of work in the colonial archives of the Dutch East Indies. Earlier generations of radical historians had urged reading archives against the grain, looking past the colonizers' perspective to recover the voices of the colonized. Stoler argued that this was necessary but insufficient. Historians also needed to read along the grain: to take the colonial archive seriously as a record of what officials worried about, what they considered knowledge, and how they reasoned. The archive was not only a source about colonized people but a document of colonial thought, and its categories, its repetitions and its anxious gaps were themselves evidence. The literary scholar Saidiya Hartman pressed the problem further. In her essay "Venus in Two Acts," published in the journal Small Axe in 2008, she confronted the archive of the Atlantic slave trade, in which enslaved women appear mostly as entries in ledgers, as victims in court cases, or as names in ship records, their lives visible only at the moments when they were being bought, sold, punished or killed. Hartman asked what a historian owes to people who survive in the record only through the violence done to them. Her answer, which she called critical fabulation, was a form of writing that acknowledged the limits of the archive while refusing to accept its silences as final. It did not invent facts, but it reasoned openly about what the record could not say. Stoler and Hartman were writing about paper. Their methods become more urgent in the digital environment, for three reasons. First, digitization often strips records of the context that allows reading along the grain. A page image in a database, retrieved by keyword search, arrives without the surrounding files, the register entries and the marginalia that told a researcher how the document fitted into the bureaucracy that produced it. Second, digital systems add their own grain, the categories of metadata schemas and the logic of search engines, which researchers must learn to read as carefully as they read colonial categories. Third, the scale of digital archives encourages methods such as keyword searching and computational analysis that can reproduce the archive's biases at scale. If the colonial archive recorded the colonized mainly as problems, then a search of that archive for mentions of a community will return a portrait of that community as a problem, now with the apparent authority of data. Why the colonial case matters for the digital present It might be objected that the story of the migrated archives belongs to an older and cruder form of archival power, in which a state deliberately destroyed and hid evidence of its own misconduct. Surely digital archives, run by libraries and platforms rather than colonial governors, are a different matter. The objection has some force. Most digital curatorial decisions are made without malice, often by engineers who have never thought about their work as archival. But three features of the colonial case recur in the digital present, and they justify beginning here. The first is the separation of record from subject. The colonial files on Kenya were kept in Britain, far from the people they described, and under British control. Much of the digitized cultural heritage of the global South is similarly held on servers in the global North, under the control of institutions and companies whose legal and commercial obligations lie elsewhere. The question of where a record physically sits, and under whose law, has not disappeared with digitization. It has become a question about data centres and jurisdictions. The second is the invisibility of selection. For decades the existence of the Hanslope Park files was not publicly acknowledged; researchers did not know what they were not seeing. Digital selection is invisible by default. A user of a digital library sees what has been digitized, not what has been left out, and nothing on the screen tells her that the collection she is searching represents a small and particular fraction of the physical holdings. The third is the gap between documentary and lived memory. The Kenyan claimants carried their history in their bodies and their communities. The courts, and for a time the historical profession, gave greater weight to the paper. In the digital environment, the same hierarchy persists in new forms. What has been recorded and indexed is findable and citable; what exists as oral tradition, community knowledge or unindexed material is increasingly treated as though it did not exist. The risk is not only that some memories are lost, but that the digital record becomes the standard against which all memory is judged. The stakes Archives matter because societies use them to settle arguments about the past, and those arguments have consequences in the present. The migrated archives mattered because they affected whether elderly torture survivors would be compensated. Records of land ownership matter in disputes over restitution. Records of human rights abuses matter in trials. Records of scientific research matter when findings are challenged. Records of ordinary life matter because they allow people to see themselves in history and to argue that their experience counts. When those records move into digital form, the stakes do not diminish. They are amplified, because digital records can be copied and circulated at a scale paper never allowed, and because they can also be deleted, corrupted or buried more quickly and more quietly. The theorists surveyed in this chapter offer a vocabulary for thinking about archival power: custody and appraisal, the archontic joining of keeping and interpreting, the four moments of silence, reading along and against the grain. The rest of this book applies that vocabulary to the specific machinery of digital memory, beginning with the act that transfers the past into the digital realm in the first place: digitization. Chapter 2: From Shelf to Server In December 2004 Google announced that it would scan the collections of several of the world's great research libraries, among them Harvard, Stanford, the University of Michigan, the New York Public Library and Oxford's Bodleian. The ambition was breathtaking. Michigan alone held millions of volumes, and Google proposed to digitize them all, making their full text searchable by anyone with an internet connection. The company's co-founders spoke of organizing the world's information. Librarians who had spent careers watching budgets shrink found themselves courted by one of the richest corporations on earth, offering to do in a decade what their institutions could not have done in a century. The project, which became Google Books, changed how scholars work and how the public finds old texts. It also provoked a decade of litigation, a political backlash in Europe and a lasting debate about what happens to cultural heritage when a private company becomes its principal gateway. Those debates are usually framed in terms of copyright. This chapter frames them differently. Digitization is not simply the transfer of existing material to a new medium. It is a new act of selection, description and transformation that creates a new object, the digital surrogate, with its own biases and its own politics. To understand the digital archive we have to understand what is gained and lost in the passage from shelf to server. Mass digitization and its discontents The Google Books project faced immediate legal challenge. The Authors Guild and a group of publishers sued in 2005, arguing that scanning in-copyright books without permission was infringement. The parties negotiated a sweeping settlement that would have created a registry for rights holders and allowed Google to sell access to out-of-print books. In 2011 Judge Denny Chin of the federal district court in New York rejected it, in part because it would have given Google a de facto monopoly over millions of so-called orphan works whose rights holders could not be found. The litigation continued. In 2013 Chin ruled that Google's scanning and display of short snippets was fair use, and in 2015 the Court of Appeals for the Second Circuit affirmed, holding that the creation of a searchable index was sufficiently transformative. The Supreme Court declined to hear the case in 2016. The legal victory secured the searchable index, but the broader vision of a universal digital library faded. Google's public statements about the project became quieter, and the full texts of in-copyright books remained locked away, visible only as snippets. What survived most visibly was HathiTrust, a partnership founded in 2008 by research libraries that had contributed volumes to Google, which pooled their digital copies into a shared repository. HathiTrust faced its own suit from the Authors Guild and won in the Second Circuit in 2014, with the court upholding full-text search and access for print-disabled readers as fair use. HathiTrust now holds millions of digitized volumes, but public-domain works are fully readable only where the law permits, and the rules differ between the United States and elsewhere. In Europe the project provoked a different kind of reaction. Jean-Noël Jeanneney, then president of the Bibliothèque nationale de France, published a short polemic in 2005 whose title can be translated as When Google Challenges Europe. He argued that allowing an American company to decide which books would be digitized, and how they would be ranked in search results, risked subordinating European culture to American commercial priorities. Jeanneney was not opposed to digitization. He was opposed to leaving it to a single private actor whose selection logic reflected its home market. His argument helped build political support for Europeana, a portal launched by the European Union in 2008 to aggregate digitized collections from museums, libraries and archives across member states. Europeana illustrates a subtler form of curatorial power. It does not itself digitize. It aggregates metadata supplied by thousands of contributing institutions, and so its coverage reflects which institutions had the resources and the will to digitize, and to describe their work in the formats Europeana requires. Wealthy national libraries in western and northern Europe contributed heavily; smaller institutions in poorer regions contributed less. The result is a portal that presents itself as European cultural heritage but reflects the uneven distribution of money and technical capacity across the continent. No one designed that imbalance. It emerged from the sum of many institutional decisions, which is exactly how Trouillot's silences usually form. Digitization as selection Every digitization programme selects. Even the most ambitious mass digitization projects covered only a fraction of the world's documentary heritage, and the fraction was not random. Printed books were favoured over manuscripts because they are easier to scan. Material in Latin script was favoured over other writing systems, because optical character recognition worked better on it. Collections in wealthy institutions were favoured because those institutions could negotiate partnerships and pay for their share of the work. Material free of copyright was favoured because it could be displayed in full, which meant that in many countries works published before the early twentieth century became far more visible online than works from the middle of the century, whose rights status was murky. The historian Lara Putnam examined the consequences in an article in the American Historical Review in 2016, "The Transnational and the Text-Searchable." She argued that digitized sources had enabled a remarkable expansion of transnational history, allowing historians to follow people, ideas and goods across borders by searching for names and terms in collections they could never have visited. But she also warned that the new methods cast shadows. Historians who could search digitized English-language newspapers from the comfort of their offices might neglect the undigitized local archives that would give them context. What she called "side-glancing," the incidental knowledge gained by sitting in an archive and reading around the document one came for, was being lost. The result could be scholarship that was broad but thin, confident about connections between places while ignorant of the places themselves. Canadian historian Ian Milligan offered harder evidence. In an article in the Canadian Historical Review in 2013, he studied Canadian history dissertations and found that after two major national newspapers were digitized and made searchable, citations to those papers rose sharply, while citations to papers that had not been digitized did not follow the same pattern. The digitized newspapers had become more important in the historical record not because they were more significant but because they were easier to use. Milligan's warning was that the convenience of digital sources was quietly reshaping historical interpretation, and that historians were rarely reflecting on the fact. Selection has a long pre-digital history. In the second half of the twentieth century many libraries microfilmed their newspaper runs and then discarded the originals, arguing that film was a more durable and space-efficient format. The novelist Nicholson Baker attacked this practice in Double Fold, published in 2001, arguing that libraries had destroyed irreplaceable artefacts on the basis of exaggerated claims about paper decay, and that the microfilm copies were often poor, incomplete and in black and white where the originals had been in colour. Baker's critics said he romanticized paper and underestimated the storage crisis libraries faced. But his central point, that a surrogate is not the original and that choosing surrogates over originals is a consequential decision, applies with equal force to digitization. When institutions treat the scan as the object, and deaccession or neglect the physical item, the scan's limitations become the limitations of the record. The hidden labour of description A digital image of a page is almost useless without description. To be findable, it needs metadata: a title, a date, a creator, subjects, and technical information about the file. It usually also needs text, extracted from the image through optical character recognition, so that its contents can be searched. Both processes are laborious, both involve judgement, and both introduce distortions that users rarely notice. Subject headings are the clearest example of how description encodes politics. The Library of Congress Subject Headings, used by libraries across the English-speaking world and beyond, long included the heading "Illegal aliens" for works about undocumented immigrants. In 2014 students at Dartmouth College, working with librarians there, petitioned the Library of Congress to change it on the grounds that the term was dehumanizing. In 2016 the Library announced it would replace the heading. The decision became a political controversy: members of Congress objected, and legislative language was introduced to prevent the change. The Library of Congress did not complete a revision until 2021, when it replaced the heading with terms including "Noncitizens" and "Illegal immigration." The episode, documented in the 2019 film Change the Subject, showed that the words used to organize knowledge are not technical details. They are contested descriptions of people, and changing them can require years of political struggle. Optical character recognition introduces a different kind of distortion. Early OCR systems were trained on modern typefaces and struggled with historical printing, especially the long s, ligatures, broken type, faded ink and multi-column newspaper layouts. The resulting text could be riddled with errors, so that a keyword search for a name or a place would miss many instances where the OCR had misread it. Because the errors were not uniform, the gaps were systematic: poorer-quality printing, smaller newspapers, non-English text and unusual typefaces were all more likely to be misread and therefore less likely to appear in search results. Researchers who relied on keyword searching without understanding OCR quality were, in effect, sampling from a biased population without knowing it. The historian Tim Hitchcock, co-director of large digitization projects including the Old Bailey Proceedings Online, reflected on these issues in an essay titled "Confronting the Digital," published in Cultural and Social History in 2013. Hitchcock argued that historians had adopted digital tools with enthusiasm but without adequate critical reflection on how those tools shaped their evidence. The keyword search in particular, he suggested, fragmented sources and encouraged historians to treat documents as bags of words rather than as artefacts produced in specific circumstances. His solution was not to reject digital methods but to demand that historians understand them as well as they understood their paper sources. The surrogate and the original The digital surrogate is not the original, and the differences matter. A scan captures the visual surface of a page at a particular resolution and in particular lighting. It usually loses the texture of the paper, the watermarks, the binding, the smell, the weight, and the physical relationship between one document and those around it. Scholars of the history of the book have long argued that these material features carry meaning: the quality of paper can indicate the wealth of a publisher or the date of a printing, and annotations and wear show how a book was used. A scan that crops the margins or flattens the binding can remove this evidence. Digitization also often separates documents from their archival context. In a physical archive, a letter sits in a folder, in a box, in a series, and its position tells a researcher something about who kept it and why. In many digital collections, individual items are presented as standalone objects, retrieved through search, with the archival hierarchy reduced to a line of metadata or omitted entirely. This is what archivists mean when they insist on provenance and original order, principles that go back to nineteenth-century European archival practice. Stripping context is not a neutral simplification. It changes what a document can be understood to mean. There is another difference, which is the relationship between surrogate and original over time. A digitized collection is often treated as finished once it goes online. But the physical originals continue to exist, and in some cases continue to be added to, re-described or reinterpreted. The digital version may freeze an old description that the institution has since revised, or reflect an arrangement that has since changed. Conversely, the physical original may be neglected or disposed of on the assumption that the digital copy is sufficient. The surrogate can become a replacement without anyone ever deciding that it should. Born-digital records and the collapse of distinctions So far this chapter has treated digitization as the transfer of paper material to digital form. But an increasing share of the record is born digital, created on computers and never existing on paper. Email, word-processed documents, websites, social media posts, databases and digital photographs make up the bulk of the record of the early twenty-first century. They raise distinct problems, because there is no physical original to fall back on. If the digital version is lost or corrupted, nothing remains. Born-digital records also collapse distinctions that traditional archives relied upon. A physical letter is a discrete object. An email exists within a server, a thread, a set of attachments and metadata about senders and recipients, and it can be displayed in many ways depending on the software used to read it. A website is a set of files assembled by a browser at the moment of viewing, and it may look different on different devices, draw content from other servers and change constantly. The question of what exactly an archive should preserve, the files, the appearance, the behaviour or the experience, has no single answer. The Library of Congress confronted these problems in its most ambitious born-digital acquisition. In 2010 it announced an agreement with Twitter to receive the entire archive of public tweets since the service began in 2006. The announcement generated excitement among researchers who imagined a comprehensive record of public conversation. But the Library struggled to make the collection usable. The volume of data was enormous, the technical infrastructure needed to provide access was expensive, and privacy and legal questions were complex. At the end of 2017 the Library announced that it would stop collecting all tweets and would instead acquire them selectively from the start of 2018, on a limited basis around events and themes of significance. The comprehensive archive of public conversation became, in the end, another selective archive. The Twitter episode is instructive because it shows how the fantasy of comprehensiveness collides with the realities of cost, law and technology. Nobody decided that most tweets were historically worthless. The decision emerged from the practical impossibility of managing the whole, which is the same pressure that led Schellenberg to theorize appraisal seventy years earlier. Digital capacity changed the scale of the problem, not its nature. What digitization changes Digitization, then, is neither the democratization of the past nor its betrayal. It is a transformation that creates new possibilities and new silences at once. It makes vast amounts of material accessible to people who could never have reached it before, enables forms of analysis impossible with paper, and allows fragile originals to be consulted without being handled. It also privileges certain kinds of material over others, strips context, introduces invisible errors, concentrates power over access in whoever controls the digital infrastructure, and risks replacing originals with surrogates whose limitations become the limitations of history itself. The appropriate response is not suspicion of digitization but literacy about it. Users of digital collections need to know what has been selected and what has not, how items have been described and by whom, what the quality of the extracted text is, and what physical and archival context has been lost. Institutions need to publish that information rather than hide it behind clean interfaces. And everyone needs to remember that the digital collection is the result of many decisions made by particular people for particular reasons, rather than the past simply made visible. The next chapter turns to a problem that digitization was supposed to solve but has in many ways made worse: the fragility of the record itself. Chapter 3: The Fragile Record In 1986, to mark nine hundred years since William the Conqueror's great survey of England, the BBC launched a project it called the Domesday Project. Schoolchildren and volunteers across the United Kingdom were invited to document their local areas in text and photographs, and the results were combined with maps, statistics and video into a multimedia portrait of the nation. The material was stored on special laser discs designed to be read by a BBC Micro computer fitted with a particular player. It was an extraordinary piece of technology for its time, and a deliberate echo of a document that had survived nearly a millennium on parchment. Within about fifteen years the new Domesday Book had become almost unreadable. The hardware was obsolete, the players were rare and failing, and the software depended on a computing environment that no longer existed. The original Domesday Book, meanwhile, sat safely in the National Archives at Kew. In the early 2000s a research project called CAMiLEON, a collaboration between the University of Leeds and the University of Michigan, developed an emulator that allowed the 1986 content to run on modern computers. Later the BBC made parts of the material available online again. But the episode became a standard warning in the digital preservation literature, cited again and again as proof that digital records are fragile in ways that parchment and paper are not. This chapter examines that fragility. It argues that the vulnerability of digital records is not a technical accident that better engineering will eventually solve. It is a structural feature of digital memory, arising from the fact that digital objects exist only as long as a chain of hardware, software, institutions and funding continues to support them. Because that chain is maintained by particular people and organizations, the question of what survives becomes a question of who keeps paying and who keeps caring. Fragility is where curatorial power operates in its most basic form: the power to let things die. The anatomy of digital decay Digital decay takes several forms, and they are often confused. The first is physical: storage media degrade. Magnetic tapes lose their magnetization, optical discs delaminate, hard drives fail mechanically and solid-state storage loses charge over time. The informal term bit rot describes the gradual corruption of stored data as individual bits flip. Unlike paper, which usually decays gradually and visibly, a digital file can go from perfect to unreadable when a small amount of corruption hits a critical part of its structure. The second form is format obsolescence. A file is only meaningful if software exists to interpret it. File formats designed by commercial companies for their own products are often poorly documented, and when the software that reads them is discontinued, the files become inaccessible even if the bits are perfectly preserved. The end of Adobe Flash is a recent example on a vast scale. For two decades Flash powered animations, games, educational materials and entire websites. Adobe ended support at the close of 2020, and browsers removed the ability to run Flash content. Thousands of games and artworks that defined a generation's experience of the web became unplayable in ordinary browsers overnight. Preservationists responded with emulation projects, including an open-source Flash emulator called Ruffle, and the Internet Archive hosts a large collection of Flash works that can be played through it. But the rescue was partial and depended on volunteers and nonprofits acting quickly. The third form is dependency failure. Many digital objects depend on other systems to function: a web page that loads images from another server, an application that requires a licensing server to start, a database that needs a particular operating system. When any of those dependencies disappears, the object breaks, sometimes in ways that are hard to detect. A web page that once displayed an embedded video may now show only a blank rectangle, and a future reader may not know that anything is missing. The fourth form is institutional abandonment. Digital objects require continuous maintenance: storage must be migrated to new media every few years, formats must be monitored for obsolescence, and systems must be patched and upgraded. All of this costs money and requires skilled staff. When an organization loses funding, changes priorities or goes out of business, its digital holdings can disappear with it. This is the most common cause of digital loss and the least technical. As Table 2 sets out, the forms of decay differ in cause but share a common remedy in sustained institutional commitment. Table 2. Principal threats to digital records and typical responses. Threat Typical example Usual response Media degradation Failing tapes, corrupted drives Replication, integrity checking Format obsolescence Flash content, proprietary files Migration, emulation Dependency failure Broken embeds, dead licence servers Capturing dependencies, web archiving Link rot Dead URLs in citations Persistent identifiers, archived snapshots Institutional abandonment Closed services, lost funding Legal deposit, distributed custody Near misses and actual losses The history of digital preservation is full of near misses that reveal how thin the margin can be. One of the most frequently told concerns the animated film Toy Story 2. During production in 1998, a command run on Pixar's servers began deleting the film's files, and the company discovered that its backup system had not been working properly. The film was saved largely because Galyn Susman, a technical director who had been working from home after the birth of her child, had a copy of the files on her home computer. A multimillion-dollar production by one of the most technically sophisticated studios in the world came within one employee's personal habits of being lost. A less happy story concerns the images taken by NASA's Lunar Orbiter missions in the 1960s. The spacecraft photographed the Moon in high resolution, and the data was recorded on analogue tapes. The images released at the time were degraded copies. The original tapes survived, but the specialized tape drives needed to read them had become almost unobtainable. In 2008 a project called the Lunar Orbiter Image Recovery Project, working in a disused fast-food restaurant at NASA's Ames Research Center in California, restored old drives and recovered images at far higher quality than had ever been seen. It was a triumph, but it depended on a handful of enthusiasts, a stockpile of surviving hardware, and the fact that someone had kept the tapes. Had any link in that chain been broken, the data would have been gone. Some losses have not been averted. In 2019 MySpace, once the dominant social network for musicians, acknowledged that a server migration had resulted in the loss of music and other content uploaded over roughly twelve years. The lost material included songs by millions of artists, many of them amateurs whose only copies of their early work had been on the platform. MySpace said little about the loss beyond a brief acknowledgment. Shortly afterwards the Internet Archive made available a substantial set of MySpace audio files that had been gathered for an academic study, which preserved a fraction of what was lost. The rest is simply gone. The MySpace case shows how a private company's routine technical operation can destroy a cultural record without any public process, any accountability or any mechanism for affected creators to intervene. Link rot and the crisis of citation A particular form of digital decay has attracted attention because it strikes at the foundations of scholarship and law. Link rot is the phenomenon by which web addresses cease to work: the page has been moved, deleted or replaced, and the link returns an error or leads somewhere else. A related problem, sometimes called reference rot or content drift, occurs when the link still works but the content at the address has changed, so that a reader following a citation finds something other than what the author cited. In 2014 Jonathan Zittrain, Kendra Albert and Lawrence Lessig published a study in the Harvard Law Review Forum examining links in legal scholarship and in opinions of the United States Supreme Court. They found that a large majority of the links cited in the law journals they studied, and about half of the links in Supreme Court opinions, no longer led to the material originally cited. The implications were serious. Legal reasoning depends on the ability to check sources, and the Supreme Court's opinions are among the most authoritative texts in the American legal system. If the evidence underlying those opinions was disappearing from the web within years, the authority of the law itself was being quietly undermined. The authors did not simply diagnose the problem. They helped build a response: Perma.cc, a service developed at Harvard Law School's library that allows authors to create permanent archived copies of web pages they cite, with a stable link that libraries commit to maintaining. Courts and law journals have adopted it widely. The episode shows that link rot is not inevitable. It results from the absence of institutions committed to maintaining links, and it can be addressed when such institutions are created. The scale of the problem across the web as a whole was measured by the Pew Research Center in a study published in 2024. Pew found that around 38 percent of web pages that had existed in 2013 were no longer accessible a decade later. It also found that many news websites and government websites contained broken links, and that more than half of the Wikipedia pages it examined had at least one broken link in their references. For a society that increasingly treats the web as its public record, these findings describe a record that is dissolving as it is written. The digital dark age In 2015 Vint Cerf, one of the engineers who designed the protocols on which the internet runs and by then a vice-president at Google, used a scientific meeting to warn of a possible digital dark age. His concern was that future historians might find the early twenty-first century harder to study than the nineteenth, because so much of its record existed in formats that would become unreadable. Cerf proposed what he called digital vellum: a system in which data would be preserved together with the software and hardware descriptions needed to interpret it, so that a future user could reconstruct the environment in which it was created. Cerf's warning received wide publicity, and it was useful in drawing attention to the problem. But the phrase digital dark age can mislead. It suggests a uniform loss, as though the whole record of an era might vanish. In practice the loss is selective. Well-funded institutions preserve their records with care. Government records in countries with strong archival laws are often transferred to national archives. Major news organizations maintain their own archives. What is most at risk is the record of the less powerful: personal websites, small community organizations, independent media, activist groups, amateur creators and the vast informal culture of the web. The digital dark age, if it comes, will not fall evenly. It will fall hardest on those whose voices were already marginal in the traditional archive. This is the point at which fragility becomes a question of power. When preservation depends on resources, those with resources are preserved. When it depends on institutional mandates, those whom institutions consider significant are preserved. When it depends on volunteers, those whom volunteers care about are preserved. Each of these mechanisms reproduces existing hierarchies of attention and value. The fragility of digital records thus amplifies the silences Trouillot described, at the moment of fact assembly, in ways that are harder to reverse than the silences of paper archives. A paper letter left in an attic may be found a century later. A website on a server whose bills stopped being paid will not be. Building institutions of permanence Digital preservation is not a lost cause. Over the past three decades the field has developed standards, tools and institutions designed to address the problems described here. The most influential standard is the Open Archival Information System reference model, developed initially by the space-data community and adopted as an international standard in 2003. It provides a framework for thinking about what a digital archive must do: receive material, store it, manage its description, plan for its long-term preservation and provide access. The model is abstract, but it gave the field a common vocabulary and helped institutions think systematically about responsibilities that had previously been handled piecemeal. Another important development was the principle of distributed custody. The LOCKSS programme, whose name stands for "Lots of Copies Keep Stuff Safe," began at Stanford University Libraries in the late 1990s under David Rosenthal and Vicky Reich. Its premise was that the safest way to preserve digital material is to keep many copies in many independent locations, with software that checks the copies against one another and repairs any that have become corrupted. The idea is old, as monasteries copying manuscripts understood, but LOCKSS turned it into a technical system that libraries could deploy together. Its underlying insight is political as well as technical: preservation is safest when it does not depend on any single institution's continued existence or goodwill. Legal deposit is a third mechanism. Many countries have long required publishers to deposit copies of printed works with a national library. Extending this obligation to digital publications, and to the web, has been one of the most important preservation developments of recent decades. The United Kingdom introduced regulations for non-print legal deposit in 2013, allowing the British Library and other deposit libraries to harvest websites published in the UK. Other national libraries have similar programmes. Legal deposit converts preservation from a voluntary activity into a public obligation, and it gives national institutions the authority to capture material without seeking permission from every publisher. Yet even these institutions are vulnerable. In October 2023 the British Library suffered a ransomware attack attributed to a criminal group called Rhysida. The attack disabled the library's online systems, including its catalogue, for months, and stolen data was offered for sale online. The library published a detailed account of the incident in 2024, acknowledging that parts of its technical infrastructure had been old and difficult to secure. The attack did not destroy the library's collections, but it demonstrated that even a national institution with a legal mandate and a long history can be knocked out by a determined attacker. The digital archive is fragile not only because of decay, but because it is a target. Fragility as a political fact It is tempting to treat digital fragility as an engineering problem awaiting an engineering solution: better storage, more robust formats, smarter emulation. Those things help. But the cases in this chapter point to a different conclusion. Almost every digital loss described here, and almost every rescue, turned on institutional decisions. MySpace lost its music because a company did not treat its users' uploads as something worth protecting. The Lunar Orbiter images survived because enthusiasts decided to save them. Link rot in legal citations was addressed when a law library decided to build an institution to prevent it. The British Library's systems were vulnerable in part because of long-term decisions about investment in infrastructure. This means that digital preservation is a question of governance. Who has the obligation to preserve? Who pays? Who decides what is worth the cost? Who is accountable when something is lost? At present the answers are scattered across public institutions with legal mandates, private companies with none, and nonprofits and volunteers who fill the gaps without any secure footing. The next chapter examines the most important of those nonprofits, the Internet Archive, and the legal and political struggles that have tested whether an independent institution can serve as the memory of the web. Hashtags: #TheDigitalArchive #DigitalMemory #CuratorialPower #DigitalErasure #ArchivalPower #DigitalPreservation #Digitization #BornDigitalRecords #AlgorithmicCuration #ArchiveNeutrality #ArchivalSilences #MichelRolphTrouillot #ArchiveFever #MetadataPolitics #OpticalCharacterRecognition #DigitalSurrogates #WebArchiving #LinkRot #FormatObsolescence #BitRot #InstitutionalAbandonment #InternetArchive #DistributedCustody #DigitalStewardship #FutureOfDigitalMemory
- The Economics of Aging (Silver Economy Markets and Gerontechnology)
Download the Book (PDF): Introduction In the spring of 2024 Oji Holdings, one of Japan's largest paper companies, announced that its subsidiary Oji Nepia would stop making baby diapers for the Japanese market that September. The company had once produced around 700 million infant diapers a year; by 2024 the figure had fallen to roughly 400 million. Japan had recorded 758,631 births in 2023, the lowest number since the nineteenth century. Oji would keep making diapers, but for adults. Its larger rival Unicharm had crossed the same line more than a decade earlier: in Japan, sales of adult incontinence products have exceeded sales of baby diapers since 2011. Stories like this are usually told as a demographic curiosity or as a warning. They are more useful as a business case. A paper company did not need a theory of aging to see that its market had moved. It needed to understand who would now buy its product, who would use it, and who would pay for it. For a baby diaper those questions have simple answers: parents choose, parents pay, and the infant has no say. For an adult incontinence product the answers are tangled. The user is an adult with strong views about dignity and discretion. The purchase may be made by a spouse, a daughter, or a nursing home procurement manager. The bill may be paid out of pocket, by a care home's operating budget, or, in Japan, partly through public long-term care arrangements. Every one of those parties wants something different from the same product. That tangle is the subject of this book. Its argument can be stated in a sentence. The markets created by population aging are three-party markets, in which the older person who uses a product or service, the person who chooses it, and the party who pays for it are often different people with different interests, and underneath all three sits a chronic scarcity of care labour. The businesses that last are those that align the user, the chooser and the payer, and that either economise on care labour or make it more productive. Most failures in the so-called silver economy come from designing for only one of the three, or from assuming that technology can remove the need for human care rather than change its shape. Why a narrow argument It is possible to write about the economics of aging as a survey: pensions, retirement communities, cruise lines, cosmetic surgery, reverse mortgages, dementia drugs, the political economy of entitlements. This book does not do that. It treats three markets in depth, because they show the three-party problem most clearly and because they are where the largest sums of private and public money are now being committed: products designed for older adults, the housing-with-care business typified by assisted living, and the use of remote monitoring and telemedicine to manage chronic disease. Gerontechnology, the application of engineering and digital technology to the needs of older people, runs through all three. Several subjects are deliberately left out or treated briefly. Pension design and retirement income are covered only insofar as they determine who can pay for care. Pharmaceuticals and longevity biotechnology, however fascinating, follow a logic of patents and clinical trials that is different from the logic of care markets. Financial products aimed at retirees are mentioned where they bear on the ability to pay for housing and care, not reviewed on their own terms. The aim is depth on a narrow front rather than a catalogue. The shape of the argument The first two chapters set out the ground. Chapter 1 examines what demographic change actually implies. The headline numbers are striking, and the United Nations now projects that people aged 65 and over will outnumber children under 18 worldwide by the late 2070s. But the commercially important facts are finer-grained: the rapid growth of the population over 80, the wide variation in health and capacity within any age band, and the uncertain relationship between longer life and longer disability. Chapter 2 turns to money. Older households hold a disproportionate share of wealth in rich countries, and the phrase "longevity economy" captures real spending power. Yet that wealth is concentrated, much of it is locked in housing, and a large middle group of older people will have too much to qualify for public help and too little to buy private care. The chapter introduces the distinction between user, chooser and payer that organises the rest of the book. Chapters 3 and 4 deal with products and technology. Chapter 3 asks why so many products designed explicitly for older people fail, and why some of the most successful products used by older people were never marketed to them at all. It draws on the history of kitchen tools, simplified phones, hearing aids, and incontinence products, and on the recent deregulation of hearing aids in the United States, which turned a medical device sold through clinics into a consumer product sold in pharmacies and, through a software update, into a feature of a mass-market earphone. Chapter 4 examines gerontechnology more broadly: emergency alarms, sensors, fall detection and care robots. It uses Japan's long experiment with robotic care to show a recurring pattern, in which technology introduced to save labour ends up creating new work because it was designed without the people who would have to operate it. Chapters 5 and 6 are about care itself. Chapter 5 works through the unit economics of assisted living, the private-pay model of housing with personal care that has grown into a large real-estate asset class in the United States. It explains why the business is so sensitive to occupancy and wages, why the separation of property ownership from care operation has produced both fortunes and failures, and how the British collapse of Southern Cross Healthcare in 2011 illustrates the dangers of that separation. Chapter 6 argues that care labour is the binding constraint on every model in the book. Paid care workers are poorly paid, hard to recruit and quick to leave; unpaid family caregivers provide the bulk of care and are valued at sums larger than entire public programmes; and countries such as Japan face projected shortfalls in the hundreds of thousands of workers. Chapter 7 examines telemedicine for chronic care, where the evidence is richer and more mixed than either enthusiasts or sceptics tend to admit. Large randomised trials of home telemonitoring have produced results ranging from reduced mortality to no effect at all, and one of the most ambitious telehealth trials ever run found that its intervention was not cost-effective at conventional thresholds. The chapter argues that what matters is not the device but the clinical service wrapped around it, and that payment rules decide which services get built. Chapter 8 generalises that point. Across Japan, Germany, England and the United States, the design of long-term care financing determines which businesses can exist. The same care need produces a thriving home-care industry in one country and an informal family economy in another, depending on who pays and how. The Conclusion draws out what follows for entrepreneurs, investors, policymakers and families, and names the questions that remain unsettled. A note on evidence The economics of aging attracts large, round, and often unverifiable numbers: trillion-dollar markets, multi-billion forecasts, compound annual growth rates stated to a decimal place. This book avoids them unless they come from an identifiable source with a transparent method. Where the evidence is weak, the text says so. Where a claim rests on a single study, the study is named. Many of the most instructive facts are not market-size estimates at all but the results of specific trials, the balance sheets of specific companies, and the rules of specific payment systems. Those are the facts that tell a business what will work. The examples come mainly from the United States, Japan, Germany and the United Kingdom, because those countries have the longest records, the best data, and the most instructive contrasts. Readers elsewhere will recognise the same forces. Aging is now a global condition. Its economics are not. Chapter 1: The Arithmetic of a Longer Life Every business plan in the silver economy begins with a demographic chart, and most of them draw the wrong conclusion from it. The chart shows the number of people over 65 rising steeply, and the inference is that a market is growing at the same rate. It is not. People over 65 do not form a market. They are a population, one that spans more than three decades of life, contains marathon runners and people confined to bed, and includes some of the richest and some of the poorest households in any society. What demographic change creates is not one large new customer but a set of shifts in the distribution of needs, capacities and resources. The commercial opportunities lie in those shifts, and they are easy to miss if one looks only at the headline. What the numbers actually say The United Nations Population Division publishes the most widely used projections of global population, revised every two years. The 2024 revision of World Population Prospects contains several milestones that are worth stating precisely, because they are often misquoted. Table 1 sets them out. Table 1. Selected global ageing milestones from the UN World Population Prospects 2024. Indicator Figure Timing Global life expectancy at birth 73.3 years, up 8.4 years since 1995 2024 Projected global life expectancy About 77.4 years 2054 People aged 80 and over 265 million, more than infants under 1 Mid-2030s Share of deaths at age 80 or over More than half, versus 17 per cent in 1995 Late 2050s People aged 65 and over 2.2 billion, more than children under 18 Late 2070s World population peak About 10.3 billion Mid-2080s Source: United Nations, World Population Prospects 2024: Summary of Results. Two features of this table matter more than the others. The first is the date at which people aged 80 and over outnumber infants: the mid-2030s, roughly a decade from now, and more than forty years before the better-known crossover between over-65s and children. The fastest-growing age group in most rich countries is not the "young old" of their sixties but the oldest old. The second is the change in the age at death. In 1995 fewer than one death in five occurred at 80 or older. By the late 2050s more than half will. A world in which most people die in their eighties and nineties is a world in which most people will spend some years living with chronic illness, frailty, or cognitive decline before they die. That is where care demand comes from. National figures sharpen the picture. Japan, the oldest large country, had more than 29 per cent of its population aged 65 or over by 2024. South Korea, Italy, Germany and much of eastern Europe are following. China, which had one of the world's youngest populations within living memory, is aging at a speed no large country has experienced before, because its fertility collapse was so abrupt. In the United States the process is slower, cushioned by higher fertility and immigration, but the arrival of the large baby boom cohorts into their eighties from the mid-2020s onward will produce a surge in the population most likely to need help with daily living. Why age is a poor proxy for need The difficulty for any business is that chronological age predicts need only loosely. Two people aged 78 may differ more from each other than either differs from a typical person of 58. One may work part-time, travel, and manage a portfolio; the other may have heart failure, diabetes, and early dementia, and require help to bathe. The variation within age groups increases as people get older. This is one of the most robust findings in gerontology, and it is one reason that products and services designed for "seniors" as a category so often miss. Gerontologists and actuaries have developed several ways to measure aging that capture this better than a birthday. One is the idea of functional status, measured by the ability to carry out activities of daily living, such as bathing, dressing, eating, toileting and moving from bed to chair, and the instrumental activities that support independent life, such as managing money, shopping, cooking and taking medication. It is limitations in these activities, not age itself, that generate demand for care. In the United States, eligibility for long-term care insurance benefits and for many Medicaid programmes is defined by the need for help with two or more activities of daily living or by cognitive impairment. Japan's long-term care insurance, discussed in Chapter 8, assigns people to care levels after an assessment of functional need. The payer, in other words, measures function rather than age, and so should anyone trying to sell into a care market. Another approach, developed by the demographers Warren Sanderson and Sergei Scherbov, is to measure age prospectively, by the number of years a person can expect to live rather than the number already lived. On that measure, a 65-year-old today is in some respects "younger" than a 65-year-old in 1970, because she has more years ahead of her. Sanderson and Scherbov showed in a 2005 paper in Nature that when aging is measured this way, some populations age far more slowly than conventional measures suggest. The point for business is that the threshold of "old" keeps moving. Products pitched at a fixed age boundary are pitched at a target that drifts away. Longer life, longer disability, or both The most important unsettled question in the economics of aging is whether the years added to life are healthy years. There are three classic positions. In 1980 the physician James Fries argued in the New England Journal of Medicine that the onset of chronic disease could be postponed faster than death, so that illness would be compressed into a shorter period at the end of life. In 1977 the psychiatrist Ernest Gruenberg had argued the opposite: that medicine's success in keeping people with chronic disease alive would lengthen the period of disability, the "failures of success." Kenneth Manton proposed a middle position, a dynamic equilibrium in which people live longer with chronic disease but in a less severe state because the disease is better managed. The evidence since then supports parts of all three. In many rich countries, the prevalence of severe disability among people in their sixties and seventies has fallen or stayed flat, which is broadly consistent with Fries. But the number of years lived with some chronic condition has risen, because conditions such as diabetes, hypertension and heart failure are diagnosed earlier and people survive them longer, which is consistent with Gruenberg and Manton. Dementia is the critical variable. Age-specific dementia rates appear to have fallen in several high-income countries over recent decades, probably because of better education and cardiovascular health, yet the number of people with dementia keeps rising because so many more people reach the ages at which it is common. For a business, this has a practical implication. The market for chronic disease management, which is the subject of Chapter 7, grows with the expansion of morbidity: more people living longer with conditions that need monitoring and adjustment. The market for personal care and housing with care, the subject of Chapter 5, grows with the number of people who lose the ability to manage daily life, which is driven above all by frailty and dementia in the late eighties and nineties. These are different markets with different growth rates, different payers, and different sensitivities to medical progress. A breakthrough that slows cognitive decline would shrink one and enlarge the other. A population of women living alone One further feature of late-life demography deserves emphasis because it shapes nearly every market in this book. Women live longer than men in almost every country, and the gap widens with age. Among people in their late eighties and nineties in rich countries, women outnumber men by a wide margin. Most women who reach very old age do so as widows, and many live alone. This has direct consequences for care. A married man who becomes frail is usually cared for first by his wife. A widowed woman who becomes frail has no spouse to care for her, and depends on children, if she has them and they live nearby, or on paid care. The typical user of assisted living, home care and personal emergency alarms is therefore an older woman living alone, and the typical family chooser is her daughter, herself often in her fifties or sixties. Women are also the great majority of paid care workers. The silver economy is, to a degree its marketing rarely acknowledges, an economy of women caring for women. It also has consequences for money. Older women on average have lower pensions and savings than men, because of lower lifetime earnings and years spent out of the paid workforce raising children or caring for parents. They must stretch those resources over longer lives. The widow in her late eighties, with a house, a modest income, and a long potential lifespan ahead, is the central figure of the middle market discussed in the next chapter, and the person for whom the gap between need and ability to pay is widest. The dependency ratio and its limits Economists often summarise population aging in a single ratio: the number of people aged 65 and over for every hundred people of working age. The ratio is useful because it indicates the pressure on pay-as-you-go pension systems and tax-financed health care, in which today's workers fund today's retirees. As it rises, governments face a choice among higher taxes, lower benefits, later retirement, or more borrowing, and every one of those choices affects the purchasing power of older people and the budgets of public payers. But the ratio overstates the problem in some respects and understates it in others. It overstates it because many people over 65 work, and the share doing so has risen in most rich countries; because working-age people are not all employed; and because productivity growth can support more dependants per worker. It understates it for the purposes of care, because the relevant ratio is not retirees per worker but people needing care per available carer. That second ratio depends on how many people are in their late eighties and nineties, and on how many working-age adults, mostly women, are available to provide paid or unpaid care. Family size has shrunk, more women are in paid work, and adult children increasingly live far from their parents. The pool of potential family carers is shrinking even as the number of people needing care rises. Chapter 6 returns to this, because it is the single most important constraint on every care business. Germany offers a concrete projection. The federal statistical office reported that about five million people needed long-term care at the end of 2021, and projected that population aging alone would raise the figure by 37 per cent, to around 6.8 million, by 2055. That projection assumes that age-specific rates of care need stay constant. If health at older ages improves, the number will be lower; if survival with dementia lengthens, it may be higher. Either way the increase is large, and it is concentrated among the oldest. Old before rich The rich countries aged slowly and grew wealthy first. France took well over a century for the share of its population over 65 to double from 7 to 14 per cent; the United States and the United Kingdom took many decades. Their pension systems, health services and care arrangements, however imperfect, were built while their populations were still relatively young, and their citizens accumulated housing and financial wealth over long working lives. The countries now aging fastest are doing so at far lower levels of income and with far less time to prepare. China is the largest example. Its fertility fell abruptly from the 1970s onward, under the pressure of development and of the one-child policy, and its population began to shrink in 2022, the first decline in some six decades. In 2024 the government legislated the first increase in its statutory retirement ages since the 1950s, to be phased in over fifteen years from 2025, raising the age for men from 60 to 63 and for women by several years. A large share of China's older people live in rural areas with thin pension coverage, and the traditional expectation that children will care for their parents collides with the migration of those children to distant cities. Thailand, Vietnam, Brazil and Iran are aging at speeds comparable to or faster than the rich countries did, at a fraction of their income. For business, "old before rich" changes the shape of demand. In such countries the market for private-pay care, of the kind examined in Chapter 5, is confined to a thin affluent urban layer. The larger opportunity lies in low-cost, scalable models: community-based services, simple technologies that work on cheap phones, and care arrangements that support rather than replace family carers. It also lies in the design of public systems. China began piloting long-term care insurance in 2016, and the choices it makes about financing will determine what kind of care industry emerges, a point Chapter 8 develops. A company that exports the business model of an American assisted living operator or a Japanese care robot to a country aging at a quarter of the income is likely to be disappointed. Why "silver tsunami" misleads Popular writing on aging reaches for metaphors of catastrophe: silver tsunami, demographic time bomb, grey wave. The metaphors are commercially misleading as well as unkind. A tsunami is a single, sudden, homogeneous event. Population aging is gradual, predictable decades in advance, and highly differentiated. The people who will be 85 in 2045 are already alive and already 65. Their health, savings, housing and family circumstances can be studied now. There is no excuse for being surprised. The metaphor also encourages a particular error, which is to treat older people as a problem to be managed rather than as customers, workers and citizens with preferences. That error shows up in products designed around deficits rather than desires, in care services designed around the convenience of the provider, and in technology designed around what engineers can build rather than what users and carers will actually do. The following chapters trace that error through several markets. The first step is to recognise that the population arriving at old age over the next thirty years is larger, older, more varied, and in many cases richer than any before it, and that it will not buy anything that tells it it is old. What the arithmetic leaves open Demography sets the scale of the opportunity but not its shape. Three questions determine the shape, and none of them is answered by population projections. The first is who has the money. A market exists only where need meets the ability to pay, and the distribution of wealth among older people is highly unequal. That is the subject of the next chapter. The second is who does the work. Most of the value in care markets is created by human labour, and the supply of that labour is shrinking relative to demand. Technology can change how that labour is used, but the record so far suggests it cannot replace it. The third is who decides. Older people are not always the decision-makers about their own care, particularly at the point where care becomes intensive. Adult children, spouses, hospital discharge planners, physicians and insurers all shape choices. A product that appeals to the user but not to the chooser, or to the chooser but not to the payer, will struggle however large the demographic tailwind. The rest of this book is an attempt to show how those three parties interact, market by market, and what that interaction means for anyone trying to build something that lasts. Chapter 2: Who Holds the Purse If older people were poor, the silver economy would be a matter for social policy alone. They are not, at least not on average. In the rich world, households headed by people over 60 hold a larger share of wealth than at any time in modern history, and their spending sustains whole industries. But averages hide almost everything that matters. The wealth of older people is concentrated in a minority of households, much of it is held in the form of a house that cannot easily be spent, and it is guarded by a fear of outliving it that makes older people cautious consumers precisely when their care costs start to rise. To understand who will pay for the products and services of an aging society, one has to look at the distribution of wealth, at the form it takes, and at the psychology of spending it down. The size of the purse The most widely quoted estimates of older people's economic weight come from a series of reports commissioned by AARP, the American association for people over 50, and prepared with Oxford Economics. Their most recent Longevity Economy Outlook estimates that adults aged 50 and over generated $12.5 trillion in economic activity in the United States in 2024, around 43 per cent of GDP, through their own spending and the activity it supports. Nearly 57 million people over 50 were in the labour force. Households led by someone over 50 accounted for about 70 per cent of charitable giving, and the same group provided unpaid caregiving and volunteer work that the report valued at $1.2 trillion. These figures are large because the group is large: "50 and over" spans half a century of life and includes people at the peak of their earnings. They are best read as a corrective to the assumption that older people are a drain rather than as a sizing of any particular market. A company selling stair lifts does not sell to the $12.5 trillion. It sells to the much smaller number of households that have a two-storey house, a member who can no longer manage the stairs, and the means and willingness to install a lift rather than move. Wealth tells a more concentrated story than spending. Data from the Federal Reserve's Distributional Financial Accounts, which reconcile household surveys with the national balance sheet, show that the baby boom generation, born between 1946 and 1964, holds roughly half of all household net worth in the United States. Similar patterns exist in the United Kingdom, where older households own most of the housing stock outright, and in Japan, where people over 60 hold the majority of household financial assets. Much commentary treats this as a vast pool of money waiting to be spent on the needs of old age, or transferred to heirs in a "great wealth transfer." The distribution within the purse The trouble with the pool is its shape. Wealth among older people is at least as unequal as wealth in the population as a whole, and on some measures more so, because inequality accumulates over a lifetime. The Federal Reserve Bank of St. Louis, analysing the same distributional accounts, found that at the end of 2024 the top fifth of households by income held about 71 per cent of all household wealth, while the bottom fifth held around 3 per cent. The same research noted that average wealth figures for any group are much higher than median figures, because they are pulled up by the richest households. When a report says that older households have average wealth of some large sum, it describes the experience of relatively few of them. This matters because the costs of care in old age are high relative to typical wealth. As Chapter 5 sets out, the national median cost of assisted living in the United States was about $74,000 a year in 2025, and a private room in a nursing home about $130,000. The median length of time people need intensive care is measured in years rather than months for those with dementia. A household with a paid-off house, a modest retirement account, and Social Security income can meet ordinary living costs comfortably and still be unable to pay for three years of residential care without selling the house. Researchers at NORC at the University of Chicago quantified this problem in a 2019 study in Health Affairs, memorably titled "The Forgotten Middle." They projected that by 2029 there would be 14.4 million middle-income Americans aged 75 and over, and that more than half of them would not have enough financial resources, even counting the equity in their homes, to pay for assisted living and their other health costs. These are households with too much income and too many assets to qualify for Medicaid, the means-tested public programme that pays for most long-term care in the United States, and too little to buy private care for long. They are the largest group of older people, and they are the group for whom the market works least well. The same structure appears elsewhere, shaped by national rules. In England, local authorities fund residential care only for people with assets below a threshold that has been fixed at £23,250 since 2010, and the value of the home is counted if nobody else lives there. Successive governments promised to cap lifetime care costs; a cap legislated in 2014 was repeatedly postponed and was finally cancelled in 2024. The result is a middle group who fund their own care until their assets fall to the threshold, often paying higher fees than local authorities do for the same bed in the same home. Chapter 5 examines why that cross-subsidy exists and what it does to the economics of care homes. A worked example makes the arithmetic concrete; the household is hypothetical but the prices are the 2025 national medians. Consider a widow of 84 who owns a house worth $350,000 outright, has $250,000 in a retirement account, and receives $28,000 a year from Social Security and a small pension. While she lives at home and manages alone, her income covers her costs and her savings grow slowly. Now suppose she develops dementia and needs to move into assisted living at the median cost of $74,400 a year, rising to memory care at a higher price as her condition progresses. Her income covers little more than a third of the base cost. The shortfall of around $46,000 a year, before any memory care premium, medical costs, or annual rent increases, would exhaust her financial savings in about five years. If she sells the house, the proceeds extend her runway by perhaps seven or eight more years. She is, by any ordinary standard, comfortably off, and she can pay for a long stay. But her neighbour with half the savings and a cheaper house cannot, and her daughter, watching the numbers, may reasonably conclude that caring for her at home for as long as possible is the only way to preserve anything at all. The arithmetic is harsher for couples. When one spouse needs residential care and the other remains at home, the household must pay for two dwellings from one set of assets. Medicaid rules in the United States include protections for the spouse who remains in the community, allowing that spouse to keep the home and a limited amount of income and assets, but the protected amounts are modest. Many couples find that the care of one spouse consumes the savings on which the other expected to live for another decade or more. Wealth that will not be spent Even where wealth is sufficient, older people often do not spend it. Economists have long been puzzled by how slowly retirees draw down their savings. The simple life-cycle model predicts that people save during working life and spend down in retirement, ending with little. In practice many retirees spend less than their income and leave substantial estates. Research by the Employee Benefit Research Institute, published in 2018, found that American retirees who entered retirement with at least $500,000 in non-housing assets had spent down only around 12 per cent of those assets after nearly two decades. Several explanations are offered, and all have implications for business. The first is uncertainty about lifespan. Nobody knows whether they will live to 75 or 100, and running out of money at 92 is a far worse outcome than leaving money unspent at death. The natural insurance against this risk is an annuity, which converts a lump sum into an income for life. The economist Menahem Yaari showed in 1965 that, under standard assumptions, people should annuitise most of their wealth. Almost nobody does. Voluntary annuity markets are small in most countries, a gap economists call the annuity puzzle. People distrust insurers, dislike losing control of capital, want to leave something to their children, and fear that an annuity will leave them unable to pay for a sudden large expense. The largest such expense is care. The second explanation for slow spending is precautionary saving against the risk of needing long-term care, which is uninsured for most people. In the United States, the private long-term care insurance market that was supposed to cover this risk has shrunk dramatically since the early 2000s. Insurers priced policies on assumptions about lapse rates, interest rates and longevity that proved wrong: fewer policyholders let their policies lapse than expected, interest rates fell, and claimants lived longer. Large insurers took heavy losses, many left the market, and those that remained imposed steep premium increases on existing policyholders. The collapse of that market means that for most middle-class Americans the only insurance against catastrophic care costs is their own savings and, at the end, Medicaid. It is rational for them to hold on to their money. The third explanation is that housing wealth is illiquid by design. For most middle-income older households the house is the largest asset, and it provides both shelter and a sense of security. Equity release products, such as reverse mortgages in the United States and lifetime mortgages in the United Kingdom, allow owners to borrow against the house without moving, but take-up has remained modest, held back by cost, complexity and, in the American case, a history of poor practice. Most people unlock their housing wealth only when they move, and most older people do not want to move until they must. When they finally do, the move is often to a care setting, which is why the timing of house sales is so closely linked to the occupancy of assisted living communities. User, chooser and payer The wealth of older people is therefore real but partial, unequal and reluctant. That is the first reason silver economy markets behave differently from ordinary consumer markets. The second reason is more structural, and it organises the rest of this book. In most consumer markets the person who uses a product also chooses it and pays for it. In markets created by aging, those three roles frequently separate. The user is the older person who wears the device, lives in the apartment, or takes the medication. Users care about dignity, comfort, independence, and not being reminded that they are old. The chooser is the person who decides what is bought. Sometimes the user chooses, but as needs become more intensive the choice passes, in whole or in part, to others: an adult child researching care options, a spouse, a hospital discharge planner who needs a bed by Friday, a physician who orders a monitoring programme, or a care home that selects the incontinence products for all its residents. Choosers care about safety, reassurance, convenience, and in the case of professionals their own workload and liability. The payer is whoever funds the purchase. It may be the user's own savings, the family's, a private insurer, a public programme such as Medicare or Medicaid in the United States, a statutory insurance fund in Germany or Japan, or a local authority in England. Payers care about cost and, increasingly, about measurable outcomes. These three parties can be arranged in many ways, and each arrangement produces a different market. Consider some examples. A personal emergency alarm worn on the wrist is usually used by an older parent, chosen by an adult child who wants reassurance, and paid for by either of them. The user often dislikes it and leaves it on the bedside table. An assisted living apartment is used by the parent, researched and in effect chosen by a daughter or son under time pressure after a hospital stay, and paid for from the parent's savings and house sale. A remote blood-pressure monitoring programme is used by the patient, chosen by the physician or health system, and paid for by Medicare. A hearing aid, until recently, was used by the patient, chosen by an audiologist, and paid for out of pocket, which is why it was so expensive and so often declined. The commercial consequences are profound. A product that delights the user but does nothing for the chooser's anxiety will not be bought. A product that reassures the chooser but embarrasses the user will be bought and then abandoned. A service that the user and chooser both want but that no payer will fund will reach only the rich. And intermediaries will emerge to profit from the gaps. In the United States, senior living referral services such as A Place for Mom are free to families and are paid by the communities into which they place residents, commonly through a fee tied to the first month's rent. The chooser gets guidance at no visible cost; the community pays for the customer; and the referral service's incentives run towards placement rather than towards the best outcome for the user. That is not a scandal. It is simply what happens when the three roles separate and nobody designs for all of them. What follows for a business Three lessons follow from the distribution and form of older people's wealth. First, the market for private-pay care and premium products is real but narrow. It is dominated by the most affluent fifth or so of older households, and competition for them is intense. Businesses that serve this segment, such as high-end senior housing, concierge medicine and luxury travel, can prosper, but they are not addressing the core of demographic demand. Second, the middle market is the great unsolved problem. It is where most need is, and it is where neither private purchasing power nor public programmes are sufficient. Companies that find ways to deliver care more cheaply, to unlock housing wealth on acceptable terms, or to plug into public or insurance funding for this group will address the largest opportunity, and the hardest one. Third, any product or service must be designed with the user, the chooser and the payer in view. The next chapter examines what happens when products are designed for only one of them, beginning with the most intimate relationship of all, between an older person and the things she uses every day. Chapter 3: Designing for People Who Refuse to Be Old In 1955 the H. J. Heinz Company, which had built a fortune on baby food, noticed that some older adults with dental problems or digestive complaints were eating its puréed jars. The company launched a line of soft foods for them under the name Senior Foods. It failed. The product was sensible, the need was real, and the manufacturing was already in place. What went wrong was the label. Few people wanted to be seen buying food for the elderly, let alone food that came in a jar reminiscent of what they had fed their grandchildren. Joseph Coughlin, who directs the AgeLab at the Massachusetts Institute of Technology, tells this story in his book The Longevity Economy as the founding parable of a recurring mistake: designing products around a narrow idea of old age as decline, and then being surprised that nobody wants to be associated with them. This chapter is about that mistake and its opposite. It argues that the most successful products used by older people are rarely marketed as products for older people, that the user's desire for dignity usually outweighs the functional advantages of a specialised product, and that the separation of user, chooser and payer described in Chapter 2 explains why so many well-intentioned designs end up in drawers. It also examines a recent natural experiment, the deregulation of hearing aids in the United States, which shows what happens when a product moves from a clinical channel with a professional chooser to a consumer channel where users choose for themselves. The stigma tax Every product that signals age or disability carries what might be called a stigma tax: a cost, paid by the user in self-image and social standing, that must be outweighed by the product's benefit before it is used. The tax is highest for products that are visible, that mark the user as dependent, and that are associated with the end of life. Walking frames, hearing aids, incontinence pads, large-button phones and emergency alarm pendants all carry a heavy tax. Many people who would benefit from them go without, or buy them and do not use them. The stigma tax explains a pattern that designers of gerontechnology meet repeatedly. A product tests well in the laboratory, where participants are recruited because they have a particular need, and then performs poorly in the market. In the laboratory, the need is legitimised and the stigma removed. At home, in front of friends and grandchildren, the product becomes a public statement about the user's decline. Emergency alarms are the classic case: the pendants that allow someone who has fallen to summon help are commonly found in drawers or on bedside tables rather than around the necks of those who fell. The tax falls unevenly on the three parties. The user pays it. The chooser, often an adult child, does not, and may not notice it; the chooser buys reassurance and is frustrated when the user refuses to wear the device. The payer, if it is an insurer or public programme, pays for a device that is not used. Products designed for the chooser's anxiety rather than the user's dignity are likely to fail in this way. Design for the young old and include everyone The alternative is to design products that do not look like products for old people at all, but that happen to work well for them. The approach has several names. In the United States it is usually called universal design, a term associated with the architect Ronald Mace, who worked from a wheelchair and led the Center for Universal Design at North Carolina State University. Mace and his colleagues set out principles in the 1990s: products should be usable by people with diverse abilities, flexible in use, simple and intuitive, tolerant of error, and require low physical effort. In the United Kingdom and in much of the technology industry the preferred term is inclusive design. A line usually credited to the British geriatrician Bernard Isaacs states the principle: design for the young and you exclude the old; design for the old and you include the young. The most famous commercial illustration is the OXO Good Grips range of kitchen tools. In the late 1980s Sam Farber, a retired housewares executive, noticed that his wife Betsey, who had arthritis in her hands, struggled with an ordinary vegetable peeler. He worked with the design firm Smart Design to create a peeler with a thick, soft, oval handle that could be gripped without strength or pinching. It was launched in 1990, together with a small range of other tools, and was sold as a better kitchen tool, not as an aid for people with arthritis. It became one of the best-known housewares brands in the world. The design served arthritic hands well because it served all hands well. Nobody paid a stigma tax to use it. The principle extends far beyond the kitchen. Kerb cuts, introduced for wheelchair users, are used by parents with pushchairs, travellers with suitcases and delivery workers with trolleys. Lever door handles, step-free showers, larger text on screens, voice assistants, and cars with higher seats all serve older users without singling them out. Automakers discovered long ago that a car marketed to older buyers struggles, while a car designed with easy entry, clear instruments and good visibility, and marketed to everyone, sells well to older buyers. The AgeLab's AGNES suit, a wearable rig that simulates the stiffness, reduced vision and loss of dexterity of an older body, was created partly to let younger designers experience these constraints and design for them without labelling the result. There is a limit to this approach. Some needs are specific and cannot be hidden inside a universal product. A person with severe hearing loss needs amplification; a person with advanced dementia needs supervision. Universal design narrows the range of products that carry a stigma tax but does not eliminate it. For the remaining products, the task is to reduce the tax through form, channel and language. The simplified phone and the chooser's market The story of the Jitterbug phone illustrates both the potential and the ceiling of products designed explicitly for older users. The telecommunications entrepreneur Arlene Harris founded GreatCall in 2005 and launched the Jitterbug, a mobile phone with large buttons, a simple menu, a loud speaker, and access to live operators who could place calls for the user. Over the following decade the company added health services, including a nurse line, medication reminders, and an urgent response button that connected users to trained agents. By 2017 GreatCall had around 800,000 subscribers and annual revenue in the region of $250 million. The private equity firm GTCR bought it that year, and in 2018 the electronics retailer Best Buy acquired it for $800 million, as the foundation of a strategy to become a provider of connected health services for older people at home. The business was later renamed Lively. GreatCall succeeded because it served a real segment: people who found smartphones confusing and valued reassurance, and adult children who wanted their parents reachable and protected. The chooser and the user were aligned, at least for a time. But the segment was shrinking. Each new cohort of older people was more comfortable with smartphones than the last, and mainstream phones added large-text modes, emergency features and voice control. A product defined by what older people could not do was racing against a population that could do more each year. Best Buy's broader ambition met a harder obstacle, which was the payer. In 2021 it bought Current Health, a remote patient monitoring company that sold to hospitals and health systems running hospital-at-home programmes. The logic was that Best Buy's installation workforce and retail brand could deliver health technology into homes. In practice the buyers were health systems whose own finances were strained and whose hospital-at-home programmes depended on temporary Medicare waivers. In 2025 Best Buy sold Current Health back to its co-founder, after recording $109 million of restructuring charges in the first quarter of that year, largely tied to its health business. Its chief executive, Corie Barry, said the at-home care business had been harder and slower to develop than expected, citing slow adoption of hospital-at-home programmes and the financial difficulties of health care providers. A consumer company had tried to sell into a market where the chooser was a hospital and the payer was Medicare, and found that its consumer strengths did not transfer. Hearing aids: a natural experiment in who chooses Hearing loss is among the most common conditions of aging and one of the most undertreated. In the United States, the Government Accountability Office reports that about 38 million adults report some hearing loss and an estimated 29 million could benefit from hearing aids. For decades only a minority of those who could benefit used them. The reasons included stigma, since hearing aids were visible and associated with old age, but also price and channel. Hearing aids were regulated as medical devices and in practice sold through audiologists and hearing aid dispensers, who fitted them as part of a bundled service. Prices of several thousand dollars for a pair were common. Traditional Medicare did not cover them. The user paid, but the professional chose, and the market was dominated by a small number of manufacturers selling through professional channels. In 2015 the President's Council of Advisors on Science and Technology recommended creating a category of hearing aids that could be sold over the counter for adults with mild to moderate hearing loss. Congress legislated for it in 2017, and the Food and Drug Administration issued a final rule in August 2022. Over-the-counter hearing aids went on sale in October 2022 in pharmacies, electronics stores and online, without a prescription or fitting. The chooser became the user. The first effects were visible quickly. Prices for over-the-counter devices started well below traditional prescription prices, new entrants from consumer electronics joined the market, and retailers began to shelve hearing aids alongside other personal electronics. Then, in September 2024, the FDA authorised the first over-the-counter hearing aid software, a feature that Apple delivered to its AirPods Pro 2 earphones through a software update. Millions of people who already owned the earphones could, after a hearing test on their phone, use them as hearing aids for mild to moderate loss. The stigma tax was close to zero, because the device was a fashionable consumer product worn by people of every age. The price of the hearing aid function, to those who already owned the earphones, was nothing. It is too early to know how much these changes will raise the share of people with hearing loss who use amplification. The GAO, reviewing the new category in 2024, noted early evidence that over-the-counter aids can be as effective as prescription aids for the target group, but found that barriers remained, including cost, the difficulty of self-assessing hearing loss, and scepticism among professionals. What the experiment has already shown is that the combination of channel, price and form shapes adoption as much as the technology itself. The stakes are more than commercial. Hearing loss in midlife and later life is associated with faster cognitive decline and a higher risk of dementia. Whether treating it slows decline is harder to establish. The ACHIEVE trial, led by Frank Lin at Johns Hopkins and published in The Lancet in 2023, randomised 977 adults aged 70 to 84 with untreated hearing loss to either a hearing intervention with hearing aids or a health education programme. Over three years, it found no difference in cognitive decline for the group as a whole. But a prespecified analysis found that the intervention appeared to slow decline among participants drawn from a long-running cardiovascular study, who were older and at higher risk, and not among healthier volunteers. The finding is suggestive rather than conclusive, but it indicates that a cheap, destigmatised hearing aid may be not just a consumer convenience but a public health intervention, which is exactly the kind of product that payers might one day be willing to fund. Incontinence and the design of discretion Incontinence returns us to the opening of this book. It affects a large share of older people, particularly women, and it carries perhaps the heaviest stigma of any common condition. For decades the products sold for it were bulky, medical in appearance, and sold in pharmacy aisles that customers were embarrassed to be seen in. The chooser in institutional settings was the care home, which bought for absorbency and cost per change rather than dignity. The consumer market changed when manufacturers designed for the user's dignity. Kimberly-Clark's Depend brand introduced products that looked and felt more like ordinary underwear and ran advertising in which younger-looking adults and celebrities dropped their trousers to show that nobody could tell. Procter & Gamble entered the market in 2014 with Always Discreet, extending a brand that women already trusted from menstrual products and deliberately placing the product in the same frame of everyday personal care rather than illness. In Japan, where the adult market overtook the infant market in 2011, manufacturers segmented products finely by level of mobility and care, from thin pads for active people to products designed to reduce the burden on carers changing bedridden patients. That last segmentation is instructive. At the active end of the spectrum, the user is the chooser and the design problem is discretion. At the dependent end, the chooser is a carer or care home and the design problem is the carer's time and the resident's skin health. The same product category contains two different markets, distinguished not by age but by who chooses. A manufacturer that understands this designs different products, sells them through different channels, and speaks to different people. Lessons for design Four principles follow from these cases. First, avoid designing around a label. Products that announce themselves as being for the old will pay a stigma tax that most users will not pay. The more successful approach is to design products that serve older users well because they serve everyone well, or to place specialised products in mainstream frames, as Apple did with hearing, and as the incontinence brands did with underwear. Second, identify who chooses. Where the user chooses, design for dignity, beauty and ease. Where a family member chooses, design for reassurance but test for the user's acceptance, since a rejected product produces neither reassurance nor revenue. Where a professional chooses, design for the professional's workflow and the institution's costs. Third, watch the channel. The hearing aid market was shaped less by technology than by the rule that professionals controlled distribution. Change the channel and the market changes. Fourth, recognise that a product defined by incapacity competes against a population that becomes more capable with each cohort. Today's 70-year-olds grew up with personal computers; tomorrow's grew up with smartphones. Products built on the assumption that older people cannot use mainstream technology have a shrinking market. Products that make mainstream technology work better for aging bodies have a growing one. The next chapter applies these lessons to the broader field of gerontechnology, where the ambitions are greater and the gap between promise and use is wider still. Hashtags: #TheEconomicsOfAging #SilverEconomy #PopulationAging #Gerontechnology #LongevityEconomy #CareEconomy #ThreePartyMarkets #UserChooserPayer #LongTermCare #CareLaborScarcity #FunctionalStatus #ActivitiesOfDailyLiving #AssistedLiving #MiddleMarketCare #AgingInPlace #UniversalDesign #InclusiveDesign #StigmaTax #HearingTechnology #OverTheCounterHearingAids #RemotePatientMonitoring #Telemedicine #CareRobotics #LongTermCareFinancing #FutureOfAgingMarkets
- The Economics of AI (A Student's Guide to Prediction Machines)
Download the Book (PDF): Introduction The greatest challenge in studying artificial intelligence in a business school is separating the science fiction from the economic reality. Ajay Agrawal, Joshua Gans and Avi Goldfarb achieved this by framing AI simply as a massive drop in the cost of prediction. However, applying classical economic models of substitutes and complements to machine learning algorithms can be an abstract and demanding exercise for students preparing for a midterm. This companion demystifies the economics of algorithms. It is written expressly to explain Prediction Machines, translating complex AI capabilities into standard microeconomic frameworks. It breaks down, step by step, how cheaper prediction changes the value of human judgment, data and action. By cutting out the technical computer science jargon, it provides the economic vocabulary needed to evaluate AI investments in an academic setting. The book and its authors Prediction Machines: The Simple Economics of Artificial Intelligence was published by Harvard Business Review Press in 2018, and an updated and expanded edition followed in 2022. Its three authors are economists at the University of Toronto's Rotman School of Management. Agrawal founded the Creative Destruction Lab, a program that helps science-based start-ups scale, and Gans and Goldfarb have been closely involved with it. That setting matters for reading the book. The authors spent years watching founders pitch machine learning ventures, and they noticed that the investors and managers in the room lacked a shared language for deciding which of these ventures made economic sense. Engineers could describe what a neural network did. Few people could say what it was worth, to whom, and why. The book's answer was deliberately modest in its tools and ambitious in its reach. The authors did not propose a new economics for a new technology. They argued that the ordinary economics taught in a first-year course is already enough, provided one asks the right question first. That question is: what has become cheap? Their answer is prediction. Everything else in the book follows from applying familiar ideas about prices, demand, substitutes and complements to that one input. The authors followed the book with Power and Prediction: The Disruptive Economics of Artificial Intelligence (2022), which shifts attention from individual decisions to whole systems of decisions and asks why the economy-wide effects of AI were arriving more slowly than the technology's capabilities would suggest. That later book is useful context, and this guide draws on it where it clarifies the earlier one, but the focus throughout is Prediction Machines itself. Why the framing matters Public discussion of AI tends to swing between two poles. At one end are claims that machines will soon think, feel and replace most human work. At the other are dismissals of AI as statistical parlour tricks. Neither pole helps a manager decide whether to fund a forecasting project, or helps a student answer an exam question about how a hospital should reorganize its radiology department. The prediction framing cuts through both. It asks the student to set aside the question of whether machines are intelligent and to ask instead what a particular system does, in economic terms. A fraud-detection model takes information the bank has and produces information the bank lacks: a probability that a transaction is fraudulent. A translation system takes a sentence in one language and fills in the most likely sentence in another. A large language model takes a string of text and fills in the most likely next piece of text. Each is a machine that converts available data into missing information. Once that is clear, the analysis becomes tractable. A fall in the cost of producing missing information will raise the quantity demanded, will lower the value of things that compete with it, and will raise the value of things that work alongside it. That last sentence contains most of the book's intellectual machinery. What competes with machine prediction? Human forecasters, rules of thumb, and the costly buffers organizations build to cope with uncertainty: inventory, slack capacity, insurance, waiting rooms. What works alongside it? Data, which feeds the machine; judgment, which says what the prediction is for and what the different outcomes are worth; and action, the ability to do something with a prediction once it arrives. The authors' central practical claim is that the winners from cheap prediction will be the owners of those complements. The controlling idea of this guide The guide is built around a single idea: AI is best understood as a fall in the price of one input, prediction, and nearly every business and social consequence follows from asking which other inputs are its substitutes and which are its complements. When students struggle with the book, it is usually because they treat each chapter as a separate insight about technology. It is more productive to treat each chapter as the same economic question asked in a different place: inside a single decision, inside a job, inside a firm's strategy, inside an economy. Holding onto that idea also makes the book easier to criticize, which is part of studying it well. If prediction is only one input among several, then the value of AI depends heavily on what happens to the others, and on whether organizations can reorganize fast enough to use predictions at all. If the framing is too narrow, it will be too narrow in identifiable ways. The final chapter takes those critiques seriously, including the argument that generative AI stretches the idea of "prediction" further than the authors originally intended. How the guide is organized The guide moves from the simple to the systemic, much as the book does. It begins with the economics of a price drop, using the historical cases of artificial light and computing to show why a cheap input does more than make existing tasks cheaper: it changes which tasks are worth doing at all. It then defines prediction precisely, explains the three kinds of data that machine prediction depends on, and shows why data has unusual economics, including diminishing returns to more of it at the level of a single model and increasing returns at the level of a business. From there the guide turns to the question every student is eventually asked on an exam: what should humans do and what should machines do? The authors answer with a four-way classification of what is known and unknown, and with the observation that humans and machines fail in different ways. That classification sets up the core of the book, its anatomy of a decision. Every decision, in the authors' account, combines a prediction with judgment and leads to an action and an outcome, with data flowing in at several points. The guide spends two full chapters here, because this is where the economics becomes most precise and where students lose most marks. Judgment is treated as the assignment of payoffs to outcomes, and expected-value reasoning shows exactly why cheaper and better prediction makes judgment more valuable, not less. A dedicated chapter then sets out the microeconomic toolkit in formal terms: demand curves, cross-price effects, derived demand and the conditions under which a new technology raises or lowers the wages of the workers who use it. Worked examples use simple hypothetical numbers, labelled as such, so the logic can be checked by hand. The later chapters widen the lens. The authors' practical tools, the AI canvas and the decomposition of workflows into tasks, are explained alongside what they imply for redesigning jobs, with a contemporary case of a company that moved quickly to automate customer service and then partly reversed course. The strategy chapter works through the book's best-known thought experiment, in which an online retailer's predictions become good enough to ship goods before customers order them, and asks when AI stops being a tool and becomes a matter for the chief executive. A second contemporary case, a property company that bet heavily on price prediction and withdrew, shows what happens when the prediction is not good enough for the business model built on it. The guide then covers the risks the authors identify, from bias to manipulation to self-reinforcing feedback, and the wider social questions of inequality, market concentration and national advantage. The final chapter brings the framework up to date. It explains why large language models fit the prediction lens more literally than many readers expect, reviews two of the most carefully designed studies of generative AI in real workplaces, and sets out the strongest objections to the book's framing. The conclusion then asks what a student should take from the book once the specific technologies it describes have changed, as they already have. Each chapter closes with key takeaways and review questions. The questions are written to test understanding of mechanisms rather than recall of examples, because that is what good examinations in this subject reward. A glossary collects the terms that recur, and a short list of further reading points to the source book and to the research that has followed it. Chapter 1: Cheap Changes Everything Economists have a habit that frustrates technologists. When a new technology arrives, they do not ask what it can do. They ask what it makes cheaper. Agrawal, Gans and Goldfarb open Prediction Machines with exactly this habit, and it is the single most important move in the book. If a student understands why the question "what became cheap?" is more useful than the question "how does it work?", the rest of the book becomes a set of applications. The logic of a price drop Start with the most basic tool in economics, the demand curve. It records a simple regularity: when the price of something falls, people buy more of it. That is true of coffee, of air travel and of electricity. It is also true of inputs used by businesses. When the price of an input falls, firms use more of it, and they use it in two distinct ways. The first way is obvious. Firms keep doing what they were already doing, but more cheaply. If a bank has always employed analysts to estimate which borrowers are likely to default, and a machine can now make those estimates for a fraction of the cost, the bank will replace some of the analysts' work with the machine's. This is the effect most public commentary focuses on, and it is real. It is also the least interesting. The second way matters more. As the price keeps falling, firms begin to use the input for tasks that were never worth doing before. At a high price, an input is reserved for its most valuable uses. At a low price, it spreads to uses that would once have seemed absurd. The authors' point is that when the price of something falls far enough, the list of things it is used for changes, and the new uses are often ones nobody planned for. A third effect follows from the first two, and it is where the economic analysis earns its keep. When the price of one input falls, the value of other goods and inputs changes too. Some become less valuable, because they were doing the same job the cheap input now does. Economists call these substitutes. Others become more valuable, because they are used together with the cheap input and more of it is now being used. Economists call these complements. The familiar textbook example is that if coffee becomes cheaper, people buy more coffee, and so they also buy more of the cream and sugar they put in it; tea, meanwhile, loses some of its appeal. The same logic, applied to prediction, generates most of the book's practical advice. Two historical price drops The authors illustrate the logic with historical cases in which the price of something fell by orders of magnitude. Two are worth understanding in some detail, because they show the three effects at work over long periods. Light For most of human history, artificial light was expensive. It came from burning wood, tallow, oil and later gas, and the amount of light a household could afford after dark was tiny. The economist William Nordhaus studied the long history of lighting and showed that the true cost of a given quantity of light has fallen by an enormous factor since pre-industrial times, far more than conventional price indices captured. His work is often used to show how ordinary statistics understate the gains from new technology. The first effect of cheap light was that people bought more of it for the same purposes: reading after dark, working longer, moving safely around the home. The second effect was that light began to be used for things that would have been unthinkable when it was dear. Factories could run night shifts. Cities could light their streets. Buildings could be designed with deep interiors far from windows, which changed architecture and the shape of offices. The third effect reached into apparently unrelated places. Cheap light reduced the value of some things, such as the premium on daylight hours and on rooms positioned for sunlight, and raised the value of others, such as the evening entertainment, education and retail activity that lit cities made possible. None of these consequences was a property of the light bulb itself. They were consequences of a price. That is the lesson the authors want students to carry forward. Arithmetic The second case is closer to home. A modern computer, the authors observe, is at bottom a machine that does arithmetic, and the relentless fall in the cost of arithmetic since the middle of the twentieth century is what the computing revolution consists of in economic terms. The three effects appear again. First, tasks that already relied on arithmetic became cheaper. Governments had long employed rooms of human "computers" to produce artillery tables, census tabulations and actuarial calculations; machines took over that work. Second, as arithmetic became cheaper still, it spread to tasks that had never been framed as arithmetic problems at all. Photography is the authors' illustration. For more than a century photography was a chemical process. Once computation was cheap enough, it became possible to represent an image as numbers and to capture, store and edit it digitally. Photography was redefined as an arithmetic problem, and the firms whose business rested on film chemistry found their core skill devalued. Music, text, maps and many other things went the same way. Third, the value of complements and substitutes shifted. The value of the human skill of doing sums by hand fell to almost nothing. The value of complements, such as software, data and the skills needed to turn a business problem into a computable one, rose sharply. The parallel the authors draw is direct. Computing turned many problems into arithmetic problems. Machine learning, they argue, is turning many problems into prediction problems. Reframing AI as prediction The authors are careful about what they are and are not claiming. The advances that attracted so much attention in the 2010s were advances in a branch of computer science called machine learning, and in particular in techniques known as deep learning. These methods did not produce general intelligence of the kind found in science fiction. What they produced was a large improvement in the quality, and a large fall in the cost, of one specific capability: taking data that is available and using it to generate information that is not. The authors call that capability prediction, and they define it broadly, as the process of filling in missing information. Chapter 2 examines that definition closely. For now, the key point is that it reaches well beyond forecasting the future. Identifying which object is in a photograph is prediction in this sense: the pixels are available and the label is missing. Translating a sentence is prediction: the source text is available and the target text is missing. Judging whether a transaction is fraudulent is prediction: the transaction details are available and the truth about the transaction is missing. Once prediction is understood this way, the three effects of a price drop can be stated for AI. · Existing prediction tasks get cheaper. Demand forecasting in retail, credit scoring in lending, and risk assessment in insurance were already prediction activities. Machine learning makes them cheaper and often more accurate. · New problems are reframed as prediction. Driving was long treated as a problem requiring a detailed set of rules for every situation. Engineers made much faster progress when they reframed it as a prediction problem: given what the sensors show, what would a good human driver do next? The authors emphasize this reframing because it shows how a price drop changes the way problems are posed, not just how they are solved. Language translation followed a similar path, moving from hand-written grammatical rules to systems that predict likely translations from large collections of translated text. · The value of other things changes. Human prediction, which is a substitute, loses value. Data, judgment and the ability to act, which are complements, gain value. The authors return to this point again and again, and much of this guide is an elaboration of it. Why "simple economics" is enough A student might reasonably ask whether the book's subtitle undersells the subject. Can AI really be understood with a demand curve and the distinction between substitutes and complements? The authors' answer is that the economics is simple in its tools but not in its applications, and that the discipline of forcing every question into those tools is what makes the analysis useful. There are three reasons the approach works. First, it separates the technology from its value. The engineering behind a model may be difficult, but a manager does not need to understand gradient descent to ask whether a better forecast would change a decision, and by how much. The economic question is always about the decision the prediction feeds into. Second, it generalizes. Because it does not depend on the details of any particular algorithm, the framework survives changes in the technology. Techniques that were state of the art when the book appeared have since been overtaken, but the question of what becomes cheap and what becomes valuable is unchanged. The final chapter of this guide tests that claim against generative AI and finds that it largely holds, with some important qualifications. Third, it produces predictions that can be checked. If AI is a fall in the cost of prediction, then the firms and workers who gain should be those who own complements, and the ones who lose should be those whose value rested on being the best available predictor. That is a claim evidence can support or refute, and later chapters look at some of that evidence. What becomes of uncertainty One further consequence of cheap prediction deserves attention at the start, because it runs through the whole book. Organizations spend a great deal on coping with uncertainty. They hold inventory in case demand turns out higher than expected. They build slack into schedules in case tasks take longer. They staff emergency rooms for the busiest plausible night. They buy insurance and write elaborate rules to handle situations they cannot foresee. Much of this is a substitute for prediction. If a retailer could predict demand perfectly, it would need far less safety stock. If a hospital could predict admissions accurately, it would need less standby capacity. The authors stress that cheaper prediction therefore reduces the value of the costly structures organizations use to buffer themselves against uncertainty. Airports offer a helpful illustration of the kind the authors favour. Travellers arrive at an airport long before their flight because they cannot predict traffic and security queues. Much of the investment in comfortable lounges, shops and restaurants inside airports is, in economic terms, a way of making that unpredictability less painful. If journey times and queues could be predicted precisely, the value of that buffer would fall. This line of reasoning is one of the most distinctive contributions of the book. It shows that the effects of AI will often appear far from the algorithm itself, in the physical layout of buildings, the size of warehouses and the rules organizations write for themselves. It also anticipates the argument of the authors' later book, Power and Prediction, that the biggest gains from AI come when organizations redesign whole systems rather than inserting a prediction into an unchanged process. A worked illustration A simple hypothetical example makes the three effects concrete. Imagine a regional grocery chain that orders fresh produce each morning. Today a buyer in each store estimates tomorrow's demand for strawberries from experience, the weather forecast and the calendar. The buyer is right often enough, but not always. When the estimate is too high, fruit spoils and is thrown away. When it is too low, shelves are empty by mid-afternoon and customers go elsewhere. Suppose the chain can now buy a demand-forecasting service that uses years of sales records, local weather, school holidays and nearby events to estimate each store's demand for each product. Consider the three effects in turn. The first effect is substitution within the existing task. The forecasting service takes over much of the estimating the buyers used to do. If the service is cheaper and more accurate, the chain uses it for strawberries and for the other perishables it was already forecasting by hand. The second effect is expansion into new tasks. Because forecasts are now cheap, the chain starts forecasting things it never bothered to forecast before: demand hour by hour so it can schedule staff more tightly, demand by aisle so it can plan shelf space, and demand for items with small margins where a human buyer's time was never worth spending. Each of these is a new use of prediction made worthwhile only by the lower price. The third effect is the revaluation of other inputs. The buyers' skill at estimating demand is worth less, because the machine now does most of it. But the chain's records of past sales, previously an accounting by-product, become an asset, because they are what the forecasting service learns from. The ability to act on the forecasts also becomes more valuable: a supplier contract flexible enough to change quantities at short notice, or a delivery system able to move produce between stores during the day, is worth more when the chain knows where demand will fall. And the judgment of managers about how much a stock-out costs relative to a kilogram of spoiled fruit becomes central, because the forecast alone does not say how much to order. Nothing in this example requires knowing how the forecasting model works. Every conclusion comes from asking what became cheap and how that changes the value of everything around it. That is the method the authors teach, and it is the method this guide applies in every later chapter. The example also shows why the effects are often slow to appear. The first effect can happen quickly, because the chain simply swaps one source of forecasts for another. The second and third effects require changes elsewhere: new staff schedules, renegotiated supplier contracts, a data system that keeps clean sales records, and managers who can state their priorities precisely enough to act on a forecast. These are organizational changes, and organizations change more slowly than software. A note on hype and caution The authors wrote at a time of intense excitement about AI, and part of their purpose was to lower the temperature. By calling AI a prediction technology, they deliberately made it sound less magical. Students should notice, however, that the framing does not make AI sound unimportant. A drop in the price of light reshaped cities. A drop in the price of arithmetic reshaped the entire economy. The authors' claim is that a drop in the price of prediction is of that general kind: a change in a basic input used across almost every industry. That claim is itself a prediction about the economy, and it can be wrong. The history of general-purpose technologies such as electricity suggests that their economic effects arrive slowly, because organizations must be redesigned before they can use the new input well. Whether the fall in the cost of prediction will produce gains comparable to those earlier technologies, and how quickly, is one of the questions the final chapter takes up. Key Takeaways · The authors' core move is to ask what AI makes cheaper; their answer is prediction. · A fall in an input's price has three effects: more use for existing tasks, new uses for tasks never framed that way, and changes in the value of substitutes and complements. · The histories of artificial light and of computing illustrate all three effects over long periods. · Many problems, including driving and translation, have been reframed as prediction problems, just as many were earlier reframed as arithmetic problems. · Cheaper prediction reduces the value of costly buffers against uncertainty, such as inventory, slack and elaborate rules. · The framework's strength is that it separates the value of AI from the details of the technology. Review Questions 1. Using a demand curve, explain the difference between the effect of cheaper prediction on existing prediction tasks and its effect on new applications. 1. Identify one substitute and one complement for machine prediction in a retail bank, and explain the expected change in the value of each. 2. Why do the authors treat the reframing of driving as a prediction problem as significant, rather than merely a technical detail? 3. Explain how a fall in the cost of prediction could reduce the value of a hospital's standby capacity. 4. What does the history of photography suggest about which firms are most exposed when an input becomes cheap? Chapter 2: Prediction and the Economics of Data If AI is a fall in the cost of prediction, then two things need to be understood precisely: what prediction is, and what it is made from. The authors answer the first question with a deliberately broad definition. They answer the second with a close look at data, the input that machine prediction cannot do without. Both answers matter for exams, because questions about AI investment usually turn on whether a proposed use really involves prediction and whether the organization has, or can get, the data that prediction requires. Prediction as filling in missing information In everyday speech, prediction means saying what will happen in the future. The authors use the word more widely. For them, prediction is the process of filling in missing information: taking information that is available, which they call data, and using it to generate information that is not. This definition has three features students should grasp. First, the missing information need not lie in the future. It can lie in the present or the past. A system that reads a medical scan and estimates whether a tumour is present is predicting something about the present that no one yet knows. A system that estimates what a damaged historical document originally said is predicting something about the past. What matters is that the information is missing to the decision maker, not that it has yet to occur. Second, prediction produces a statement about likelihood, not certainty. A fraud model does not say that a transaction is fraudulent; it says the transaction has, for example, a certain probability of being fraudulent. This is why prediction on its own never settles what to do. A probability must be combined with an assessment of what is at stake before anyone can act, which is the subject of Chapters 4 and 5. Third, the definition lets the authors see prediction in places where people do not usually look for it. When a machine recognizes a spoken word, it predicts which word the sound corresponds to. When it labels a photograph, it predicts which object a human would say the image shows. When it drives a car, it predicts what a skilled human driver would do given what the sensors show. In each case, the problem has been recast so that its core step is filling in something unknown from something known. How machine prediction differs from older statistics Statistical prediction is not new. Economists, actuaries and epidemiologists have used regression analysis for more than a century. The authors explain what is different about machine learning in terms that avoid technical detail. Traditional regression requires the analyst to specify in advance which variables matter and roughly how they relate to the outcome. The analyst might propose that a borrower's default risk depends on income, existing debt and employment history, and the regression estimates how much each factor matters. The method works well when the analyst knows what to look for and when the relationships are reasonably simple. Machine learning methods ask much less of the analyst. Given a large amount of data, they search for patterns among many variables, including interactions between them that no analyst would think to specify. They can work with inputs such as images, sound and text, where the relevant variables are not obvious at all. The trade-off is that they need a great deal of data, and the patterns they find can be hard to interpret. A regression tells the analyst why it predicts what it predicts. A deep learning model often cannot, at least not in a form a person can easily follow. The authors' economic point is that machine learning made prediction cheaper and better in exactly the situations where older methods struggled: complex environments with many variables and rich data. That is why the drop in the cost of prediction has been so uneven. Where data is plentiful and the environment is stable, machines predict very well. Where data is scarce or the world keeps changing, they do not. Three kinds of data The authors distinguish three roles data plays in machine prediction, and the distinction is one of the most useful in the book for evaluating a proposed AI project. The same piece of information may play more than one role, but the roles are different and each has its own economics, as Table 1 summarizes. Table 1. The three roles of data in machine prediction. Role What it does When it is needed Example in lending Training data Builds the initial model Before deployment Past loans and whether they were repaid Input data Feeds the model to produce a prediction Every time a prediction is made A new applicant's income and history Feedback data Improves the model after deployment Continuously, as outcomes arrive Whether newly approved loans are repaid Training data is the historical record from which a model learns. To predict whether a loan will be repaid, a model needs many past loans together with information about whether each was repaid. Training data is typically the most expensive to assemble, because it must be collected, cleaned and, often, labelled. A model that recognizes diseases in images needs images that experts have already classified. Input data is what the model uses to make a particular prediction once it is running. For the lending model, it is the information on a new application. Input data must be available at the moment the decision is made, which is a practical constraint students often overlook. A model trained on information that is only known after a loan is issued is useless for deciding whether to issue it. Feedback data is information about how predictions turned out, which is used to improve the model over time. When approved loans are later repaid or not, that outcome can be fed back into the model. Feedback turns a static tool into a learning one, and the authors argue that this is where much of the long-run value lies. A model that improves as it is used grows more valuable over time, and the organization that runs it at scale accumulates an advantage that is hard for others to copy. The three roles raise different questions for a manager. For training data: do we have enough historical examples, and are they representative of the cases we will face? For input data: will the information the model needs be available, reliable and legally usable at the moment of decision? For feedback data: will we learn how our predictions turned out, and how quickly? A lending model gets feedback slowly, because loans take years to mature. A model that predicts which advertisement a user will click gets feedback in seconds. The feedback problem Feedback data has a subtle complication that the authors highlight and that recurs in the chapter on risk. An organization only observes the outcomes of the actions it actually took. A bank that rejects an applicant never learns whether that applicant would have repaid. If the model's early predictions were biased against some group, the bank will collect little feedback about that group and may never discover the error. Economists call this a selection problem. It means that feedback data is shaped by past decisions, and a model trained on it can entrench those decisions rather than correct them. One remedy is experimentation. An organization can occasionally act against the model's recommendation, approving a few applicants the model would reject, in order to learn what happens. This is costly in the short run, which is why the authors present it as a trade-off between exploiting what the model knows and exploring to learn more. Technology companies that run large numbers of controlled experiments on their users are, in this sense, paying to generate feedback data. Labels, proxies and the question of what is being predicted A further issue sits underneath all three roles of data, and it is the one most often missed in practice. A model can only predict what its training data records. Very often the outcome an organization cares about is not recorded directly, so the model is trained on a proxy: something that is measured and that is believed to track the real objective. Consider a hypothetical employer that wants to predict which job applicants will become excellent employees. "Excellent" is not a column in any database. The employer has, perhaps, past performance ratings given by managers, records of who was promoted, and records of who stayed with the firm for more than two years. Each is a proxy. Each captures something about excellence and something else as well. Performance ratings reflect the preferences and blind spots of the managers who gave them. Promotion reflects the opportunities that happened to be available. Tenure reflects pay, location and family circumstances as much as ability. A model trained to predict any of these will predict that proxy well, and it will carry the proxy's distortions into its recommendations. The economic lesson is that the choice of what to predict is not a technical detail. It is a judgment about what the organization values, expressed through the choice of label. The authors' framework makes this visible by insisting that prediction be linked to a decision and to the payoffs that decision produces. If the label does not match the payoff, the prediction can be highly accurate and still lead to poor decisions. Chapter 9 returns to this problem, because several of the best-documented cases of algorithmic bias arose from exactly this gap between the proxy a model learned and the objective its users cared about. Regulation adds a further constraint. Data protection laws, such as the European Union's General Data Protection Regulation, which took effect in 2018, restrict how personal data may be collected and used, and they vary from country to country. For a manager, this means that the supply of data is partly determined by law, and that a model which is technically feasible may not be lawful to train or run. In the language of this chapter, regulation raises the cost of some data and so lowers the returns to the predictions that depend on it. The economics of data Data is the most important complement to machine prediction, and so, by the logic of Chapter 1, its value rises as prediction becomes cheaper. But data has unusual economic properties, and the authors examine two of them in particular. Diminishing returns to data At the level of a single prediction task, more data generally improves accuracy, but each additional observation adds less than the one before. A model trained on a thousand examples improves a great deal when given ten thousand; it improves much less when moved from a million to ten million. This pattern of diminishing returns is familiar from basic production theory: adding more of one input while holding others fixed eventually yields smaller and smaller gains. Diminishing returns have practical consequences. They mean a smaller competitor with a reasonable amount of data may be able to build a model almost as accurate as a large incumbent's. They also mean that the value of data depends on what it covers. Additional observations of situations the model already handles well are worth little. Observations of rare or unusual situations, where the model is weak, can be worth a great deal. A self-driving system gains little from another thousand hours of motorway driving in clear weather and much from a handful of examples of unusual hazards. Increasing returns to scale at the level of the business The authors then make a point that appears to contradict the first, and students should be able to explain why it does not. While the returns to data diminish for a single model's statistical accuracy, the economic returns to being slightly more accurate than competitors can increase. In many markets, a small advantage in prediction quality translates into a large advantage in customers. If one search engine returns slightly better results than another, most users will choose it. Those additional users generate more data, which improves the model, which attracts more users. This is a feedback loop with increasing returns, and it can lead to markets dominated by a few firms. The logic is similar to network effects, where a product becomes more valuable as more people use it, though the mechanism runs through learning rather than through direct connections between users. The authors use this reasoning to explain why large technology firms have invested heavily in data and why a lead in AI can be self-reinforcing. Chapter 9 returns to the consequences for market concentration. The two properties fit together once one distinguishes the statistical question from the economic one. Statistically, each additional observation matters less. Economically, what matters is relative accuracy in a market where customers reward the best. A small statistical edge can therefore be worth a large economic prize. Data as an asset If data is valuable, organizations must decide how to acquire it. The authors discuss several strategic choices. A firm can collect data through its own operations, which requires deploying products that generate it. It can buy data from others, which works only if the data is available and the seller does not become a competitor. It can offer free or cheap services in return for data, which is the model behind many consumer internet businesses. Each choice has costs. Collecting data may require launching a product before it is good enough, accepting early mistakes in return for the feedback that will improve it. Buying data may leave a firm dependent on a supplier. Offering services in exchange for data raises questions of privacy and trust. The authors argue that these are strategic questions, not technical ones, and that the right choice depends on how much the data is worth to the firm's predictions and how unique it is. Data that any competitor can buy confers no lasting advantage. Data that only one firm can generate, because only that firm sees a particular set of transactions or users, can be the foundation of a durable position. The authors add a warning. Not all data is equally useful, and more data is not always the answer. Data that does not match the decision at hand, or that reflects past practices the organization wants to change, can mislead a model. A manager evaluating an AI project should therefore ask not only how much data exists but whether it is the right data: relevant, representative, available at the moment of decision and legally usable. Why this matters for evaluating AI investments The concepts in this chapter give students a checklist for any proposed AI application. Is there a clear piece of missing information that, if filled in, would improve a decision? Is there training data that relates available information to that missing information? Will input data be available when the decision is made? Will the organization learn from outcomes, and how fast? Is the data unique enough to create an advantage, or will competitors have the same? A project that cannot answer these questions is unlikely to create value, however impressive the underlying technology. A project that can answer them well has the basic ingredients for success, though, as the following chapters explain, prediction is only part of what a good decision requires. Key Takeaways · The authors define prediction as filling in missing information, which includes the present and past as well as the future. · Machine learning differs from traditional statistics mainly in needing less specification from the analyst and more data, and it performs best in data-rich, stable environments. · Data plays three roles: training, input and feedback. Each raises distinct practical questions. · Feedback data is shaped by past decisions, which creates a selection problem that experimentation can partly address. · Returns to data diminish for a single model's accuracy, but small accuracy advantages can produce increasing economic returns and market dominance. · Data is a strategic asset whose value depends on relevance and uniqueness, not only quantity. Review Questions 1. Give an example of a prediction in the authors' sense that concerns the present rather than the future, and identify the available and missing information. 1. A retailer has ten years of sales records but cannot observe prices charged by competitors at the moment it sets its own. Which role of data is the constraint, and why does it matter? 2. Explain why a bank that uses a model to reject loan applicants may never discover that the model is biased. 3. Reconcile the claims that returns to data are diminishing and that data can generate increasing returns to scale. 4. Under what conditions is data unlikely to provide a lasting competitive advantage? Chapter 3: The Division of Labour Between Humans and Machines Once prediction is recognized as an input that machines can now supply cheaply, the obvious question is who should do the predicting. The popular framing sets humans against machines, as if the only issue were which is better. The authors reject that framing. Humans and machines are good at different kinds of prediction and fail in different ways, and the economically interesting question is how to combine them. Adam Smith's idea of the division of labour, in which output rises when tasks are allocated to those best suited to them, is the right starting point. Where human prediction falls short The authors begin with an uncomfortable fact: in many settings, human prediction is poor. A large body of research in psychology and behavioural economics, much of it associated with Daniel Kahneman and Amos Tversky, documents systematic errors in human judgment under uncertainty. People give too much weight to vivid recent events. They see patterns in random sequences. They are overconfident in their own forecasts. They are swayed by irrelevant information, such as the order in which options are presented, and they are inconsistent, giving different answers to the same question on different occasions. The authors use professional baseball as a well-known illustration. Michael Lewis's book Moneyball describes how the Oakland Athletics, a team with a small budget, found value by relying on statistical analysis of player performance rather than on the intuitions of experienced scouts. The scouts were paying attention to features such as a player's physique and style that looked impressive but predicted little about how many runs the player would help produce. The statistics captured what actually mattered. The story is useful because it shows expert humans being systematically wrong in a way that a more disciplined approach to data could correct. A more consequential example comes from research on bail decisions by the economists Jon Kleinberg, Sendhil Mullainathan and colleagues, published in the Quarterly Journal of Economics in 2018. Judges deciding whether to release a defendant before trial are, in effect, making a prediction about whether the defendant will flee or commit another offence. The researchers trained a machine learning model on a large set of past cases and found that, in their simulations, decisions guided by the model's predictions could have reduced crime substantially without jailing more people, or jailed substantially fewer people without increasing crime. The judges appeared to be responding to signals in the courtroom that did not help them predict, and to be inconsistent across similar cases. The study also had to confront the selection problem described in Chapter 2, since outcomes are only observed for defendants who were released, and much of the paper's technical effort went into dealing with it. Machines have several advantages here. They are consistent: the same input produces the same output. They do not tire. They can process many more variables than a human can hold in mind, and they can detect complex interactions among those variables. Above all, where there is plenty of data about a recurring situation, they can learn the statistical regularities more precisely than a person can. Where machine prediction falls short Machines have their own weaknesses, and the authors are equally clear about these. Machine learning is only as good as the data it learns from, and it performs poorly when data is scarce or when the future differs from the past. Three limitations stand out. First, machines struggle with rare events. If a situation has occurred only a few times in the historical record, there is little for a model to learn from. Humans, by contrast, can often reason about rare situations by analogy, drawing on experience from other domains. An experienced executive who has never faced a particular kind of crisis may still judge it better than a model that has seen it twice. Second, machines struggle with causation. A model learns associations in the data. It does not, on its own, know which way causation runs or whether an association would survive a change in behaviour. This matters most when the decision maker wants to predict the effect of an action the data was not generated by. Third, machines can be fooled. Because models respond to patterns in input data, someone who understands those patterns can sometimes manipulate the input to produce the prediction they want. This is taken up in Chapter 9. Four kinds of knowledge To organize these strengths and weaknesses, the authors borrow a well-known classification of knowledge into known knowns, known unknowns, unknown unknowns and unknown knowns, and apply it to prediction. The classification is summarized in Table 2 and deserves careful study, because it is a common source of exam questions. Table 2. Four prediction situations and who handles them better. Situation What characterizes it Better predictor Illustration Known knowns Rich data on a recurring situation Machine Credit card fraud screening Known unknowns Little data, but the gap is recognized Human, often A rare industrial accident Unknown unknowns The event has never been anticipated Neither A novel kind of crisis Unknown knowns Model is confident but misled by hidden causes Human judgment needed Price and demand moving together Known knowns These are situations where data is plentiful and the relationship between available information and the missing information is stable. Card fraud, product recommendations, spam filtering and demand forecasting for established products are typical. Machines excel here, and in the authors' view this is where the drop in the cost of prediction has its most immediate effect. Human predictors in these areas face the strongest substitution. Known unknowns These are situations where the decision maker recognizes that information is missing, but there is too little data for a machine to predict well. Rare events are the classic case. The authors note that humans are sometimes surprisingly good at predicting from very few examples, because they bring general knowledge and analogies to bear. For these situations, human prediction remains valuable, although it is not always reliable, and machines can still help by organizing the scarce data that exists. Unknown unknowns These are events that no one anticipated and for which there is no relevant data at all. The writer Nassim Nicholas Taleb popularized the term "black swan" for such events. Neither humans nor machines can predict them, because there is no pattern to learn from. The appropriate response is not better prediction but robustness: building organizations that can survive surprises. The authors are candid that this is a hard limit on what prediction machines can do. Unknown knowns This is the most subtle category and the one students most often misunderstand. It covers situations where a model produces a confident prediction that is wrong, because the data hides the true causal structure. The model "knows" something, but what it knows is misleading. The standard illustration, which the authors use in a form similar to this, concerns hotel prices. In historical data, hotel room prices and occupancy tend to rise together. Prices are high during busy periods because the hotel is responding to strong demand. A naive model trained on this data might conclude that raising prices leads to higher occupancy. Any hotel manager knows that is backwards: raising prices on a quiet night will not fill rooms. The association in the data reflects a hidden cause, the level of underlying demand, which drives both variables. Economists call this problem endogeneity, and it is exactly what econometrics has spent decades learning to address. The lesson is that human understanding of how the world works, including an understanding of what causes what, remains essential when predictions are used to guide actions that change the situation. A model trained on data generated by one set of behaviours may give poor guidance once behaviour changes. The authors suggest that people with training in causal inference, including economists, have a particular role in identifying these situations. Combining human and machine prediction If humans and machines have different strengths, the natural next step is to combine them. The authors describe several ways of doing this. The simplest is to let each handle the cases it is best at. The authors call one version of this prediction by exception. A machine handles routine cases, where it is confident and data is rich, and flags unusual or uncertain cases for human attention. Fraud screening often works this way: the model clears the vast majority of transactions automatically and routes a small number of ambiguous ones to human investigators. The machine's predictions of its own uncertainty become an input to the allocation of human effort. A second approach uses the machine's prediction as an input to a human decision, or the human's assessment as an input to the machine. The study of breast cancer pathology by Dayong Wang and colleagues at Harvard, released in 2016, is a well-known demonstration. A deep learning system and a human pathologist each examined images of lymph node tissue for signs of metastatic cancer. The pathologist alone performed better than the system alone. But when the system's predictions were combined with the pathologist's diagnoses, the researchers reported that the human error rate fell by about 85 percent. The two made different kinds of mistakes, so the combination was more accurate than either. A third approach reverses the usual direction and uses the machine to check the human. A model can flag cases where a human's decision departs sharply from what the data suggests, prompting a second look. This can reduce inconsistency without removing human discretion. Combining humans and machines is not automatically beneficial. Humans may defer too readily to a machine, accepting its predictions even when they have good reason to doubt them. Economists and psychologists sometimes call this automation bias. Humans may also override a good machine too often, especially if they distrust it or if its errors are more visible than their own. Designing the division of labour well requires attention to these behavioural responses, and later research on generative AI, discussed in Chapter 10, found both patterns at work. Comparative advantage, not absolute advantage Students trained in international trade theory will recognize a familiar trap in the question "is the machine better than the human?" David Ricardo showed two centuries ago that a country can gain from trade even if its partner is better at producing everything, because what matters is comparative advantage: which party gives up less to produce each good. The same logic applies to the division of labour between humans and machines, and it clarifies several points the authors make. First, a machine does not need to be more accurate than a human to take over a prediction task. It needs to be cheap enough relative to its accuracy. Consider a hypothetical insurer that reviews claims for signs of fraud. A skilled human reviewer catches, say, nine out of ten fraudulent claims but can review only a few dozen claims a day. A model catches eight out of ten but can review every claim the insurer receives, at trivial cost per claim. The model is less accurate on any single claim, yet the insurer will almost certainly use it on the bulk of claims, because reviewing every claim at slightly lower accuracy catches far more fraud in total than reviewing a small sample at higher accuracy. The human reviewer's time is then reallocated to the claims the model flags as uncertain or high in value. Second, even where a machine is better than a human at every prediction, it may still pay to keep humans on some of them, because human time spent on prediction has an opportunity cost that depends on what else the human could be doing. As machine prediction gets cheaper, the opportunity cost of using human time for prediction rises relative to its alternatives, and humans shift towards tasks that machines cannot do at all, notably judgment and action. Third, comparative advantage changes over time. As models improve and data accumulates, the set of known-known situations expands, and tasks move from human to machine. What looked like a stable division of labour in one year may shift in the next. This is why the authors encourage managers to watch for thresholds at which a model becomes good enough to change the allocation, a theme developed as the tipping point in Chapter 8. Machines that predict people A final wrinkle deserves mention because it shows how intertwined human and machine prediction have become. Some of the most successful machine prediction systems work by predicting what a human would do. The authors describe how progress in autonomous driving accelerated when engineers trained systems on data from human drivers, so that the machine learned to predict the action a skilled driver would take in a given situation. The machine is not reasoning about traffic law. It is predicting human behaviour. This has two implications. It means human actions are a form of training data, and so the people whose behaviour is recorded are supplying a valuable input, often without being paid for it. It also means that such systems inherit the strengths and weaknesses of the humans they imitate. A model that predicts what an average driver would do will be about as good as an average driver in common situations, and will have no special insight into situations that average drivers handle badly. Getting beyond human performance requires feedback about outcomes, not only observation of human actions. Implications for the value of human skills Return to the economic logic of Chapter 1. Machine prediction is a substitute for human prediction in known-known situations. The value of human skill in those situations should fall. A person whose job consisted largely of recognizing patterns in abundant data faces strong competition from machines. But the division of labour also reveals complements. Human skill at handling rare events, at recognizing when a model's confident prediction rests on a misleading association, and at deciding what to do with a prediction all become more valuable as machine prediction spreads. The authors argue that the most valuable human capabilities in a world of cheap prediction are those machines cannot supply: understanding causes, reasoning from few examples, and, above all, judgment, which is the subject of the next two chapters. This analysis also warns against a common error in thinking about jobs. A job is rarely a single prediction task. It is a bundle of tasks, some involving prediction and many involving other things. Even where machines take over the prediction part, the other parts remain, and they may become more important. Chapter 7 develops this point. Key Takeaways · Human prediction is subject to systematic biases and inconsistency; machines are consistent and excel where data is rich. · Machines struggle with rare events, with distinguishing causation from association, and with deliberate manipulation. · The four-way classification assigns known knowns to machines, known unknowns often to humans, unknown unknowns to neither, and unknown knowns to human causal judgment. · Unknown knowns arise when hidden causes create misleading associations, as when price and demand rise together. · Prediction by exception and other combinations can outperform either humans or machines alone, but behavioural responses such as automation bias complicate them. · Cheap machine prediction lowers the value of human pattern recognition in data-rich settings and raises the value of causal reasoning and judgment. Review Questions 1. Explain, using the hotel pricing example, why a highly accurate model can give misleading guidance about the effect of an action. 1. Why can neither humans nor machines predict unknown unknowns, and what should an organization do about them? 2. Describe how prediction by exception allocates work between humans and machines in fraud detection. 3. The combined human and machine error in the pathology study was much lower than either alone. What condition must hold for combination to help in this way? 4. Using the language of substitutes and complements, explain how the spread of machine prediction affects the value of different human skills. Hashtags: #TheEconomicsOfAI #PredictionMachines #ArtificialIntelligenceEconomics #CheapPrediction #MachineLearning #PredictionAsInformation #SubstitutesAndComplements #HumanJudgment #DecisionMaking #AIInvestment #TrainingData #InputData #FeedbackData #DataEconomics #DiminishingReturnsToData #IncreasingReturnsToScale #MachinePrediction #HumanMachineCollaboration #ComparativeAdvantage #PredictionByException #AutomationBias #AlgorithmicDecisionMaking #OrganizationalRedesign #GenerativeAI #FutureOfAIEconomics
- The Economics of Immigration (Labor Markets, Remittances, and Assimilation)
Download the Book (PDF): Introduction A construction worker in Guatemala City and a construction worker in Houston can have the same training, the same stamina, the same ability to read a blueprint and pour a slab. The one in Houston will earn several times as much. Nothing about the worker explains the gap. The gap is a property of the place: of the capital, institutions, infrastructure, legal order and density of other productive people that surround a job and multiply what one pair of hands can produce. Move the worker across the border and most of the gap closes within months. Keep the worker out and it stays open for a lifetime. That single fact is the starting point for the economics of immigration, and it is the fact most public argument skips. Debates about immigration tend to begin with the receiving country and stay there. Does immigration lower wages? Does it strain public budgets? Does it change the culture of the neighbourhood? These are legitimate questions and this booklet treats each of them seriously. But they are questions about how a gain is divided, and they only make sense once the size of that gain is understood. International migration is, by a wide margin, the most powerful tool yet discovered for raising the incomes of poor people. The losses and costs that dominate headlines are, in most of the evidence we have, small by comparison and concentrated in particular places and groups. The argument of this booklet can be stated in a sentence. Migration produces a large economic surplus, most of it captured by migrants themselves; its effects on the wages of native workers and on public finances are, on average, modest and often positive, but they are unevenly distributed; and the effects on countries of origin, through the return flow of money, skills and ideas, are considerably better than the familiar story of brain drain suggests. The real economic question is therefore not whether migration makes the world richer. It does. The question is who shares in the gain, who bears the concentrated costs, and whether the institutions that govern migration are designed to spread one and cushion the other. Why the numbers matter Roughly 304 million people were living outside their country of birth in 2024, according to the United Nations, about 3.7 per cent of the world's population. That share has risen only slowly over half a century. Most people never migrate, and most who do move to a neighbouring country rather than across an ocean. What has changed more dramatically is the money that migrants send home. Officially recorded remittances to low- and middle-income countries reached an estimated 685 billion US dollars in 2024, on World Bank figures. That is more than those countries received in foreign direct investment, and several times what they received in official development aid. In Tajikistan, remittances were equivalent to about 45 per cent of national output. In Tonga, Nicaragua, Lebanon and Samoa they exceeded a quarter. These numbers alone should unsettle the habit of treating immigration as a purely domestic policy matter. Decisions made in Washington about visa categories, in Berlin about recognition of foreign qualifications, or in Riyadh about whether a worker may change employers are, in their effects, among the largest development policies in the world. A small change in the number of people allowed to move from a poor country to a rich one can do more for the incomes of that poor country's families than any aid programme likely to be designed for it. At the same time, the concentrated effects are real. A town where a meatpacking plant recruits hundreds of foreign workers in a single year faces pressure on its schools and housing that a national average conceals. A group of workers whose skills closely match those of new arrivals may see their wages grow more slowly than they otherwise would. An earlier cohort of immigrants often competes most directly with the next. A country that trains doctors at public expense and watches a third of them leave has grounds for complaint, even if the aggregate effect of emigration on its development is positive. A serious account of the economics has to hold both the aggregate gain and the local strain in view at once. What economists actually know The economics of immigration has been transformed over the past three decades by better data and by a methodological shift towards natural experiments. Early studies compared cities with many immigrants to cities with few, and struggled with the obvious problem that immigrants choose to go where jobs are plentiful. More recent work exploits sudden, unplanned shocks: a boatlift from Cuba, a policy that let Czech workers commute into German border towns, the dispersal of refugees across Danish municipalities by administrative lottery, a visa lottery that let some Tongans move to New Zealand and not others. These designs have not settled every argument, and some of the fiercest disputes in the field concern the reanalysis of the same episode by different scholars. But they have narrowed the range of plausible answers considerably. The results are, in broad terms, reassuring for the receiving country and emphatic for the migrant. Average wage effects on native workers are close to zero over the medium run. Where negative effects appear, they tend to fall on workers who are the closest substitutes for newcomers, frequently earlier immigrants. The fiscal contribution of an immigrant depends heavily on age at arrival, education and legal status, and on whether the costs of educating their children are counted as a cost of immigration or as an investment in the next generation of taxpayers. Emigration of skilled workers from poor countries often raises, rather than lowers, the stock of skills at home, because the prospect of leaving motivates people to train. And the children of immigrants, across a very wide range of countries of origin and eras, climb the income ladder faster than the children of natives who started from the same rung. None of this means that immigration is costless or that every policy choice is benign. It means the costs are of a particular kind: distributional, local and manageable, rather than aggregate and ruinous. That distinction matters enormously for policy. A cost that is aggregate calls for restriction. A cost that is distributional calls for redistribution, investment and better design. The shape of this booklet The first chapter sets out the scale of the gain from migration and why it exists: the place premium, the evidence from migration lotteries, and the logic of selection that determines who moves. The second and third chapters turn to the receiving country's labour market. Chapter 2 examines the central empirical dispute, over whether and how much immigration lowers native wages, through the natural experiments that have shaped it. Chapter 3 explains the mechanisms that allow labour markets to absorb newcomers with less disruption than a simple supply-and-demand diagram predicts, and identifies precisely where the losses do concentrate. Chapter 4 addresses the public purse. Fiscal accounting of immigration is unusually sensitive to assumptions, and the chapter explains why credible studies can reach different conclusions about the same population, and what the more careful ones agree on. Chapters 5 and 6 turn to the countries migrants leave. Chapter 5 reconsiders brain drain in light of evidence that emigration can create skills as well as remove them. Chapter 6 examines remittances: their scale, their behaviour in crises, their effects on households and economies, and the costs and taxes that erode them. Chapter 7 addresses assimilation, the economic trajectory of immigrants and their children after arrival, and what history and contemporary data show about how quickly gaps close. Chapter 8 draws the threads together around the question the rest of the book keeps returning to: how the surplus from migration is divided, why the politics of immigration so often diverge from its economics, and which policy designs have managed to share the gain more widely. Throughout, the booklet concentrates on the economic evidence and on the wealthy destination countries where most of that evidence has been gathered, chiefly the United States, the United Kingdom, Western Europe, Canada, Australia and the Gulf states. It does not attempt to settle the moral questions about national membership and the right to exclude, which economics alone cannot resolve. What it tries to do is make sure that whatever position a reader takes on those questions is informed by an accurate picture of what migration actually does. Chapter 1: The Place Premium Economics usually explains differences in earnings by differences in workers. People with more education, experience or scarce skills earn more; people with fewer earn less. This framework does well within a country. It fails almost completely across them. The largest single determinant of what a worker earns in the world today is not their schooling, their occupation, their gender or their age. It is the country in which they are permitted to work. Measuring the gap for identical workers The obvious objection to raw international wage comparisons is that they compare different people. The average worker in Norway has more years of schooling than the average worker in Nigeria, works with more machinery, and may differ in countless unmeasured ways. A fair comparison needs to hold the worker constant and change only the place. The most careful attempt to do this is a study by Michael Clemens, Claudio Montenegro and Lant Pritchett, published in the Review of Economics and Statistics in 2019 under the title "The Place Premium." The authors used nationally representative household surveys from 42 developing countries and matched them against US census data on immigrants born in those same countries. They compared men of the same country of birth, the same age, the same years of schooling, and in the core comparison men who had completed their education in the country of origin before moving, so that the American wage could not be attributed to American schooling. The resulting wage ratios were very large. For a typical developing country in the sample, an observably identical low-skill man earned several times more working in the United States than at home, even after adjusting for differences in the cost of living. The obvious remaining worry is selection on unobservables. Perhaps the people who migrate are unusually able, ambitious or healthy in ways surveys do not record, and would have earned more at home too. The authors addressed this directly, using the theory of migrant selection and evidence on how strongly migrants differ from non-migrants to bound how much of the gap such hidden differences could plausibly explain. Even under conservative assumptions, most of the gap survived. They concluded that the average price equivalent of migration barriers for low-skill men was greater than 13,700 US dollars per worker per year, measured at purchasing power parity. That is a lower bound, and it is a figure per person, per year, for as long as the barrier stands. To put the number in perspective, it exceeds the annual income of most people in most of the countries studied. No education programme, microfinance scheme or agricultural extension project yet evaluated produces an income gain of that magnitude for the individuals it reaches. Development economists spend careers searching for interventions that raise incomes by ten or twenty per cent. The place premium is measured in hundreds of per cent. The lottery evidence Bounds and selection models are persuasive, but the cleanest evidence comes from situations in which chance, rather than choice, determines who moves. The best known comes from the Pacific. Under New Zealand's Pacific Access Category, a fixed number of Tongan citizens each year are admitted as residents, chosen by random ballot from among those who apply and meet basic criteria such as a job offer. Because the winners and losers of the ballot are, by construction, statistically similar, comparing their later incomes gives a clean estimate of the effect of migration itself. David McKenzie, John Gibson and Steven Stillman exploited exactly this design in a study published in the Journal of the European Economic Association in 2010. They estimated that migrating raised the income of those who moved by 263 per cent, within the first year. They also showed something methodologically important. When they applied the standard non-experimental techniques that researchers use when no lottery is available, comparing migrants to similar-looking non-migrants, those techniques overstated the gain by between 9 and 82 per cent, because migrants are positively selected on characteristics that surveys miss. The selection concern is real. But even after it is removed, the gain from moving is enormous. Similar evidence comes from other quasi-random admission systems, including seasonal work schemes in Australia and New Zealand, and from studies of the United States diversity visa lottery. The details vary, but the conclusion does not. When otherwise similar people are sorted by chance into those who can work in a rich country and those who cannot, the ones who can work there earn multiples of what the others earn. Why the premium exists If the same worker can produce several times more in one place than another, the productivity must come from the place. Economists attribute it to several reinforcing sources. The first is capital. A worker in a rich country operates with far more physical capital: better tools, machinery, vehicles, buildings and energy. A farm labourer driving a combine harvester can harvest in an hour what a labourer with a sickle harvests in weeks. The worker's own effort is similar; the output is not. The second is technology and organisation. Firms in rich economies are, on average, better managed, more specialised and more tightly integrated into supply chains. Studies of management practices across countries have found large and persistent differences in the quality of management between firms in rich and poor economies, and those differences are associated with large differences in productivity. A worker joining a well-run firm acquires the benefit of its organisation. The third is institutions: the rule of law, secure contracts, reliable courts, low corruption, stable prices, functioning infrastructure. These are what allow specialisation and long-term investment to happen at all. They are expensive to build, slow to change and impossible to export in a container. The fourth is agglomeration. Productive people are more productive near other productive people. Dense markets for specialised skills, suppliers and ideas raise the output of everyone within them. This is why the place premium exists within countries as well as between them: a worker in a thriving metropolitan area earns more than an equivalent worker in a depressed region of the same country, and moving between them raises earnings. A striking implication follows. If the productivity of a worker is largely a property of the place, then blocking a worker's movement does not merely deny that individual a higher wage. It keeps a unit of labour where it produces less than it could. The world as a whole is poorer as a result. The same logic inside borders It helps to notice that the place premium is not a peculiarity of international migration. It operates within countries too, and the history of internal migration shows what happens when people are allowed to follow it. Between roughly 1910 and 1970, some six million Black Americans left the rural South for cities in the North and West, in what became known as the Great Migration. They moved from sharecropping and agricultural labour, under the legal and economic oppression of the Jim Crow South, to industrial jobs in Chicago, Detroit, Cleveland, New York and Los Angeles. Economic historians who have followed individual migrants through linked census records find that the move raised their earnings substantially relative to those who stayed. The migrants were the same people, with the same skills, working in a different place. China's economic transformation since the late 1970s offers a larger example still. The country's household registration system, the hukou, historically tied people to their place of birth and restricted access to urban jobs, housing and public services. As those restrictions were relaxed, hundreds of millions of rural workers moved to coastal cities to work in factories, construction and services. China's National Bureau of Statistics counts the stock of rural migrant workers in the hundreds of millions. Economists studying Chinese growth attribute a significant share of the rise in national productivity to this reallocation of labour from low-productivity farms to high-productivity urban employment. It is plausibly the largest movement of people in human history, and it took place almost entirely within one country's borders. Two features of these internal cases matter for what follows. The first is that the gains were large and were captured largely by the movers, as they are in international migration. The second is that the hukou system never fully disappeared: migrant workers in Chinese cities have long faced restricted access to local schools, health care and social insurance, so that many children of migrants were left behind in villages with grandparents. A barrier that does not stop movement can still shape who captures its benefits and who bears its costs. The same is true of international migration regimes, as later chapters show. The trillion-dollar estimate How much poorer? A line of research beginning with Bob Hamilton and John Whalley in 1984 has tried to estimate the global output gain from removing barriers to labour mobility. These are simulations rather than measurements, and they rest on strong assumptions about how many people would move and how productive they would become. Their results nonetheless point consistently in one direction. In a 2011 survey in the Journal of Economic Perspectives, provocatively titled "Economics and Emigration: Trillion-Dollar Bills on the Sidewalk?", Clemens gathered the available estimates and found that most implied that eliminating migration barriers would raise world output by somewhere between about 50 and 150 per cent of world GDP. By contrast, the estimated gains from removing all remaining barriers to trade in goods, or to international flows of capital, amounted to a few per cent of world output at most. No serious economist believes that all migration barriers could or should be removed overnight, and the political, social and fiscal constraints discussed in later chapters are real. The point of these estimates is comparative. Of all the distortions in the global economy, restrictions on the movement of people are by an order of magnitude the most costly. Even a modest relaxation, allowing a few per cent of the workforce of poor countries to work in rich ones, would on these calculations generate gains larger than the complete liberalisation of trade. It is fair to ask why, if the gain is so large, it is not realised through other channels. In principle, capital could flow to where labour is cheap, or goods made by cheap labour could be traded, equalising wages without anyone moving. In practice both channels work only partially. Robert Lucas noted in 1990 that capital does not flow from rich to poor countries on anything like the scale simple theory predicts, a puzzle generally attributed to the institutional and infrastructural shortcomings that also depress wages. And trade in goods cannot move the things that most raise productivity: working courts, dense cities and well-managed firms. Some services cannot be traded at all. A care worker, a builder or a nurse must be physically present. Migration remains the channel through which the place premium is most directly captured. Who moves, and who does not If the gain is so large, why do so few people move? About 3.7 per cent of the world's population lives outside their country of birth, and that share has changed little in decades. Part of the answer is legal restriction, which is the dominant constraint for people in poor countries seeking to work in rich ones. But even where movement is legally free, such as within the European Union or between states of the United States, migration rates are lower than the size of income differences would predict. Moving is costly, and not only in money. The direct costs can be substantial. A Nepali or Bangladeshi worker recruited for construction work in the Gulf often pays a recruitment fee amounting to many months of expected earnings, frequently financed by loans. An irregular journey from Central America to the United States involves payments to smugglers that can run to many thousands of dollars, along with serious physical risk. The indirect costs include separation from family, loss of social networks, language barriers, the recognition of qualifications, and uncertainty about whether a job will materialise. These costs shape who migrates. George Borjas, drawing on a model of occupational choice developed by the economist A. D. Roy, argued in a 1987 paper in the American Economic Review that the skill composition of migrants depends on the relative returns to skill in origin and destination. If the destination rewards skill more steeply than the origin, the most skilled will be most attracted to move; if the origin has steeper returns to skill, as in highly unequal countries, emigrants may be drawn from lower in the skill distribution. Later work, notably by Daniel Chiquiar and Gordon Hanson on Mexican migrants to the United States, found that the reality was often in between: Mexican migrants were drawn disproportionately from the middle of the skill distribution rather than the top or bottom, in part because the costs of moving put migration out of reach for the poorest. That last point has an important consequence for development policy. Migration from poor countries tends to rise, not fall, as those countries grow richer, at least up to a point. The poorest households cannot afford to move. As incomes rise, more families can pay the costs, finance the journey, and send a member abroad. Studies of emigration rates across countries find a hump-shaped relationship: emigration rises with income through the lower and middle ranges and falls only when a country reaches upper-middle income levels. This is sometimes called the migration hump. It means that aid programmes designed to reduce migration by raising incomes in countries of origin are, over any realistic horizon, more likely to increase it. The migrant captures most of the gain It is worth pausing on who receives the gain from migration, because it frames everything that follows. When a worker moves from a low-productivity place to a high-productivity one, the increase in output is shared among several parties. The largest share goes to the migrant, in the form of higher wages. Some goes to the migrant's family at home, through remittances. Some goes to employers and owners of capital in the destination, who gain from a larger workforce. Some goes to consumers, through lower prices for the goods and services migrants help produce. Some goes to governments, through taxes. And some native workers may gain or lose depending on whether they complement or compete with the newcomer. Most estimates suggest that the migrant's own share is overwhelmingly the largest. The gain to natives of the destination country, taken as a whole, is positive but small relative to the gain to migrants. This asymmetry explains a great deal about the politics of immigration. The people who benefit most are, by definition, not yet citizens and have no vote. The natives who gain do so diffusely, as consumers or shareholders, and often without noticing. The natives who lose, if any, tend to be concentrated in particular sectors and places, and they notice. The rest of this booklet takes that asymmetry seriously. The following chapters look closely at the effects on native workers, public budgets, origin countries and immigrants' own trajectories after arrival. In each case the question is less whether migration creates value than how that value is divided and what happens to the people on the wrong side of the division. Barriers as a price One final conceptual point clarifies much of what follows. Migration barriers can be thought of as a price, or more precisely as a tax, levied on the movement of labour. The place premium measures how high that tax is. Like any tax, it generates revenue for somebody and imposes a deadweight loss on the economy as a whole. The revenue, in this case, often goes not to governments but to smugglers, labour recruiters, visa brokers and employers who hold workers on restrictive sponsorship terms. The Nepali worker who pays a recruitment fee equal to a year's wages is paying the migration tax to a private intermediary. The farm worker whose visa ties them to one employer pays it in lower wages and weaker bargaining power. This framing matters because it suggests that the design of migration systems, not only their overall generosity, determines who captures the surplus. A system that admits the same number of people but lets them change employers, pays their recruitment costs, or charges a visa fee that flows to the public purse distributes the gain very differently from one that does not. Those design choices recur throughout the chapters that follow, and the final chapter returns to them directly. The place premium is the engine of everything else in this subject. It is why people move despite great cost and risk, why remittances are so large, why the prospect of emigration changes how people in poor countries invest in education, and why the political pressure around borders is so intense. Every other question about the economics of immigration is, in the end, a question about what happens around that engine. Chapter 2: Native Wages: What the Evidence Shows The most intuitive argument against immigration is also the simplest piece of economics most people know. If the supply of something rises and demand stays the same, its price falls. Labour is bought and sold; immigration raises the supply of labour; therefore immigration lowers wages. The argument is taught in the first week of any economics course, and it is not wrong as far as it goes. The question is how far it goes, and the answer turns out to depend heavily on what else changes when immigrants arrive. This chapter examines the evidence directly. The next explains why that evidence looks the way it does. The two should be read together, because the empirical findings are hard to believe without an account of the mechanisms, and the mechanisms are easy to dismiss as theory without the evidence. The problem of finding a comparison Estimating the effect of immigration on wages sounds straightforward. Find places that received many immigrants and places that received few, and compare what happened to wages. The difficulty is that immigrants do not arrive at random. They move to places where jobs are plentiful and wages are rising. A naive comparison will therefore tend to find that immigration is associated with higher wages, not because immigrants raise wages but because rising wages attract immigrants. Early studies in the 1980s and 1990s, using this spatial approach, generally found small effects, and critics argued that the smallness was an artefact of this bias. There was a second problem. Even if immigrants arrived at random, native workers and capital might respond by moving. If immigration depressed wages in Los Angeles, some native workers who would otherwise have moved there might go to Phoenix instead, and some who were in Los Angeles might leave. The effect of immigration would then be spread across the whole national labour market, and a comparison of Los Angeles with Phoenix would understate it. George Borjas, Richard Freeman and Lawrence Katz raised this concern in the 1990s, and it has shaped the debate ever since. Economists responded to these difficulties in two ways. One was to look for natural experiments: episodes in which a large number of immigrants arrived suddenly, for reasons unrelated to local labour demand, so that their destination could be treated as essentially random. The other was to abandon local comparisons entirely and study the national labour market, dividing workers into groups by education and experience and asking whether the groups that received more immigrants saw slower wage growth. Both approaches have produced important results, and they have sometimes pointed in different directions. The Mariel boatlift The most famous natural experiment in the field began in April 1980, when Fidel Castro unexpectedly announced that Cubans wishing to leave could do so from the port of Mariel. Over the following six months about 125,000 Cubans arrived in the United States, and roughly half of them settled permanently in Miami. The Miami labour force grew by around 7 per cent almost overnight, and the new arrivals were disproportionately workers with little formal education. It was hard to imagine a cleaner test of the supply-and-demand prediction. David Card examined the episode in a 1990 paper in the Industrial and Labor Relations Review. He compared wage and unemployment trends in Miami with those in four comparison cities with broadly similar economic histories, Atlanta, Houston, Los Angeles and Tampa. He found no discernible effect of the boatlift on the wages or unemployment of less-skilled workers in Miami, including less-skilled Black workers and earlier Cuban immigrants, the groups one would expect to be most exposed. Miami absorbed a sudden 7 per cent increase in its labour force with little visible effect on wages. The paper became a landmark, and for a quarter of a century it was a central piece of evidence that immigration effects were small. In 2017, Borjas published a reanalysis in the same journal that reached a very different conclusion. He argued that the right comparison group was not all less-skilled workers but the workers most similar to the Marielitos: non-Hispanic men of prime working age who had not completed high school. Focusing on this narrower group, he found that their wages in Miami fell sharply relative to comparison cities after 1980, by somewhere between 10 and 30 per cent. The two findings could not both be right, and the dispute prompted a wave of further work. Giovanni Peri and Vasil Yasenov applied a synthetic control method, which builds a comparison city from a weighted blend of other cities chosen to match Miami's pre-1980 trends, and found no significant wage effect for less-skilled workers. Michael Clemens and Jennifer Hunt then identified a more fundamental problem with the narrow sample. The Current Population Survey, the main data source, happened to change its methods around 1980 in a way that substantially increased the share of low-wage Black men among the small group of Miami high-school dropouts it sampled. Because that group earned less on average for reasons unrelated to the boatlift, its sudden appearance in the sample could produce an apparent wage decline even if no individual worker's wage had fallen. The sample sizes involved were very small, in some years only a few dozen workers. The Mariel debate has not ended, but it has become a lesson in the fragility of estimates drawn from small subgroups. Most economists who have examined the evidence conclude that it does not support large negative wage effects. It does illustrate a more general point: the closer one looks at a narrowly defined group, the more one's results depend on sampling noise and on the choices made in defining the group. Other natural experiments Mariel is not the only episode, and the others provide a more robust picture. Table 1 summarises several of the most influential. Table 1. Selected natural experiments on immigration and native labour-market outcomes. Episode Study Nature of the shock Main finding for natives Mariel boatlift, Miami, 1980 Card (1990); Borjas (2017); Peri and Yasenov (2019) About 7 per cent rise in Miami labour force Disputed; most reanalyses find no significant wage effect Soviet emigration to Israel, 1990s Friedberg (2001) Population rose about 12 per cent in four years No adverse effect on native wages once occupational sorting is accounted for Refugee dispersal, Denmark, 1986 to 1998 Foged and Peri (2016) Refugees allocated to municipalities by administrators Less-skilled natives moved into more complex jobs; wages and employment rose Czech commuters, German border, 1990s Dustmann, Schönberg and Stuhler (2017) Sudden right of Czech workers to take jobs across the border Local native employment fell, mainly via fewer new hires; small wage effects Syrian refugees, Turkey, 2012 onwards Del Carpio and Wagner (2015) Large, concentrated refugee inflows in border provinces Natives displaced from informal jobs; gains in formal employment for some Bracero exclusion, United States, 1964 Clemens, Lewis and Postel (2018) Removal of hundreds of thousands of Mexican farm workers No rise in wages or employment of domestic farm workers; farms mechanised Source: the studies named, as published in the Industrial and Labor Relations Review, Quarterly Journal of Economics, American Economic Journal: Applied Economics, Journal of Human Resources, American Economic Review and World Bank working papers. The Israeli case, studied by Rachel Friedberg in the Quarterly Journal of Economics in 2001, is particularly telling because of its scale. Following the collapse of emigration restrictions in the Soviet Union, Israel's population rose by around 12 per cent in four years. The immigrants were highly educated but initially worked in occupations well below their qualifications. Friedberg found that occupations receiving more immigrants did not experience slower wage growth for natives, once the tendency of immigrants to enter occupations where wages were already stagnating was taken into account. The Danish case, studied by Mette Foged and Giovanni Peri, exploited a policy under which refugees arriving between 1986 and 1998 were allocated to municipalities by the government, with little regard for local labour conditions and no choice by the refugees. Following individual native workers over time through administrative records, the authors found that less-skilled natives in municipalities receiving more refugees did not lose out. They tended to move into occupations involving more complex, less manual tasks, and their wages and employment rose relative to comparable workers elsewhere. The influx of refugees into manual work pushed natives up the occupational ladder. When effects do appear The German-Czech case shows that immigration can have real effects on specific groups. After the fall of the Iron Curtain, a policy allowed Czech workers to take jobs in German municipalities near the border while continuing to live in the Czech Republic. Christian Dustmann, Uta Schönberg and Jan Stuhler studied the resulting inflow, which was large in affected municipalities and largely unrelated to local demand. They found that native employment in those areas fell. The decline came mainly through a reduced inflow of new native workers into affected areas rather than from incumbents losing their jobs, and wage effects were small. The effects that did appear fell most heavily on older and less-skilled workers. Turkey's experience with Syrian refugees points in a similar direction. By the late 2010s Turkey hosted more Syrian refugees than any other country, concentrated in its southeastern provinces. Studies found that native Turkish workers in the informal sector, where the refugees could most readily work, were displaced in significant numbers. At the same time, some natives, particularly men, moved into formal employment, where refugees were initially unable to compete. The net effect varied sharply by education and gender: less-educated women and informal workers bore most of the cost. The United Kingdom's experience after the 2004 enlargement of the European Union, when citizens of eight Central and Eastern European states gained the right to work in Britain, has been studied intensively. Christian Dustmann, Tommaso Frattini and Ian Preston, writing in the Review of Economic Studies in 2013, found that immigration in the period they studied was associated with small wage losses for workers at the bottom of the wage distribution, below roughly the twentieth percentile, and small gains higher up. The Migration Advisory Committee, the UK government's independent adviser, concluded in a 2018 review of European migration that its effects on overall employment and wages were small, with some evidence of modest negative effects on the lowest paid. Opening a border: the Swiss case Most natural experiments involve the arrival of relatively less-skilled workers, because refugee flows and labour shocks at the lower end of the market are the ones most often sudden and unplanned. Switzerland provides a rare case of a sudden liberalisation affecting workers across the skill distribution. Under the agreement on the free movement of persons with the European Union, Switzerland progressively removed restrictions on cross-border commuters from neighbouring EU countries in the early 2000s, first abolishing the requirement that employers give priority to Swiss residents, and then eliminating limits on where in Switzerland commuters could work. Regions close to the border were exposed far more than regions further away, because commuting is only practical over short distances. Andreas Beerli, Jan Ruffner, Michael Siegenthaler and Giovanni Peri studied the effects in a 2021 paper in the American Economic Review. The liberalisation produced a large increase in the number of cross-border workers in the border regions, many of them highly educated. The authors found no evidence that the inflow reduced the wages or employment of Swiss residents on average. The wages of highly educated natives in the border regions rose relative to those further away. Firms in the exposed regions expanded, increased their productivity and innovated more, including by raising their patenting and research activity, and new firms entered the market. The easier availability of skilled labour appears to have relieved constraints on growth that firms had faced when they could hire only locally. The Swiss case is a useful complement to Mariel. It shows that the effects of a large, sudden labour-supply shock depend heavily on who arrives and on whether the firms employing them can expand. Where the newcomers bring skills that local firms are short of, the arrival can raise the productivity of the whole local economy. The national approach Borjas's most influential contribution came not from Mariel but from the national approach. In a 2003 paper in the Quarterly Journal of Economics, he divided US workers into cells by education and years of labour-market experience, arguing that workers with the same education but different experience are imperfect substitutes, and that immigration affects cells unevenly. Using census data from 1960 to 2000, he found that a 10 per cent increase in the supply of workers in a cell reduced wages in that cell by around 3 to 4 per cent. Applied to the actual immigration of 1980 to 2000, his estimates implied a wage reduction of roughly 3 per cent for the average native worker and around 9 per cent for native workers without a high school diploma. Gianmarco Ottaviano and Giovanni Peri challenged these results in a 2012 paper in the Journal of the European Economic Association. Their central point was that immigrants and natives with the same education and experience are not perfect substitutes: they tend to do different tasks, and immigrants cluster in particular occupations. Allowing for imperfect substitution, and for the adjustment of capital investment over time, they found that the long-run effect of immigration on native wages was close to zero or slightly positive. The largest negative effects fell on earlier immigrants, who were the closest substitutes for newcomers. The difference between these conclusions is largely a difference in assumptions about how substitutable various groups of workers are, and those assumptions are hard to test decisively. What both approaches agree on is that the group most exposed to competition from new immigrants is previous immigrants, and that any negative effects are concentrated among workers with the least education. The consensus position In 2017 the National Academies of Sciences, Engineering, and Medicine in the United States published a comprehensive review of the economic and fiscal consequences of immigration, edited by Francine Blau and Christopher Mackie and prepared by a panel that included economists from across the spectrum of views, Borjas among them. Its conclusions have been widely cited as the closest thing the field has to a consensus. The panel found that, measured over a period of ten years or more, the impact of immigration on the wages of native-born workers overall was very small. Where negative effects appeared, they were most likely to fall on prior immigrants and on native-born workers without a high school diploma, especially over shorter periods. The panel found little evidence that immigration significantly affected the overall employment levels of native-born workers, though there was some evidence that recent immigrants reduced the employment rate of native-born teenagers. It also found that high-skilled immigration had positive effects on the wages of many native workers, a point the next chapter develops. The US Congressional Budget Office, in its 2024 assessment of the immigration surge that began in 2021, reached broadly consistent projections. It estimated that the surge would slow wage growth for other workers slightly in the early years, with the effect reversing later as productivity gains accumulated, and that average wages would be modestly higher in 2034 than they would otherwise have been. The weight of the evidence therefore points to a clear conclusion. The simple supply-and-demand prediction of large wage losses is not borne out. Average effects on native wages are small over the medium run. Negative effects exist but are concentrated: on the lowest-skilled, on earlier immigrants, on informal workers, and in the short run before other adjustments take hold. These concentrated losses are real and matter for policy. But they are a small fraction of the gains that the same migration generates, which leaves room for policy to compensate the losers while still leaving everyone else better off. Why the simple prediction fails is the subject of the next chapter. The short answer is that labour markets are not a single pool of identical workers competing for a fixed number of jobs. When immigrants arrive, the number of jobs, the kinds of jobs natives do, the amount of capital and the technology firms use all change too. Chapter 3: How Labor Markets Absorb Newcomers The evidence of the previous chapter presents a puzzle. A sudden 7 per cent increase in Miami's labour force, a 12 per cent increase in Israel's population, the arrival of hundreds of thousands of refugees in Danish towns: in each case the simplest economic model predicts that wages should fall, and in each case the measured effects are small or absent. Either the evidence is wrong or the simple model is missing something. The evidence has survived decades of scrutiny. The model is what needs revising. What it misses is that immigration does not change only the supply of labour. It changes demand for labour, the mix of tasks natives perform, the amount of capital firms install, the technologies they choose, and the prices of goods and services. Each of these adjustments offsets part of the pressure on wages. Understanding them is also the key to understanding where the costs of immigration do fall, because the adjustments are uneven and some groups are left without them. Immigrants are also customers The most basic error in the simple model is known as the lump of labour fallacy: the belief that there is a fixed quantity of work to be done, so that every job taken by an immigrant is a job taken from a native. In reality, immigrants earn wages and spend them. They rent apartments, buy food, use transport, pay for haircuts and phone plans, and send their children to school. Each of these purchases creates demand for labour. An economy with more people in it has more work to do. This is not a minor adjustment. In the long run, a larger population with the same mix of skills and the same capital per worker should, to a first approximation, have the same wages. The United States grew from about 76 million people in 1900 to over 330 million in the early 2020s, much of the increase driven directly or indirectly by immigration, and real wages rose many times over. Population size is not what determines wages. Productivity is. The simple model is therefore relevant chiefly to the short run and to the question of composition: whether immigration changes the mix of skills in a way that makes some groups of workers relatively more abundant, and whether capital and technology take time to catch up. Different workers, different tasks The second adjustment is that immigrants and natives, even when they have the same formal education, tend to do different work. Giovanni Peri and Chad Sparber documented this in a 2009 paper in the American Economic Journal: Applied Economics. Using detailed occupational data on the tasks each job requires, they found that less-educated immigrants in the United States specialised in occupations intensive in manual and physical tasks, while less-educated natives specialised in occupations intensive in communication and language tasks. The reason is comparative advantage. A native English speaker has an edge in jobs that involve dealing with customers, coordinating with colleagues, or handling paperwork; a recent immigrant with limited English has a relative edge in jobs where language matters less. When immigrants arrive, this specialisation deepens. On a construction site, more immigrant labourers mean more demand for supervisors, estimators, inspectors and people who deal with clients and suppliers, and these roles tend to go to natives. In a restaurant, more immigrant kitchen staff support more front-of-house positions. In a hospital, more immigrant care assistants allow nurses to spend more time on work that requires their training. Immigration pushes natives up the task ladder into work that is typically better paid and less physically punishing. The Danish study described in the previous chapter captured exactly this process at the level of individual workers. Natives in municipalities receiving more refugees moved into jobs involving more complex tasks and saw their wages rise. A similar pattern appears in Italian data studied by Cristina Cattaneo, Carlo Fiorio and Giovanni Peri, who found that immigration was associated with natives moving into occupations requiring more complex skills. The mechanism is not universal, and it works less well for natives who lack the language or social skills to move into communication-intensive work, but it is a consistent feature of the evidence. Capital catches up The third adjustment is investment. When the supply of labour rises, the return to capital rises too, because each machine or building now has more workers available to use it. Firms respond by investing more. Over time, capital per worker returns towards its previous level, and with it the wage. How long this takes is an empirical question, and it is one reason why estimates of immigration's effects differ by time horizon. In the short run, before firms have had time to build new capacity, the effects on wages are more likely to be negative. Over a decade, most of the adjustment has typically occurred. This is why the National Academies panel emphasised effects over ten years or more, and why the Congressional Budget Office projected a small early slowdown in wage growth followed by a reversal. A sudden, large inflow into a place where capital cannot quickly expand, such as a border town with few firms, will produce larger short-run effects than a steady flow into a growing metropolitan economy. Technology choice A fourth adjustment operates through technology. The availability of labour affects which production methods firms choose. Where labour is scarce and expensive, firms invest in machines that replace workers; where it is abundant, they choose labour-intensive methods. Immigration therefore affects not only how much capital is installed but what kind. Ethan Lewis showed this in a 2011 study in the Quarterly Journal of Economics of US manufacturing plants. Plants in areas that received larger inflows of less-skilled immigrants during the 1980s and 1990s adopted automation technologies more slowly than similar plants elsewhere. Firms used the available labour instead of machines. This helps explain why less-skilled wages did not fall as much as a fixed-technology model would predict: immigration shifted technology in a direction that increased demand for exactly the kind of labour immigrants supplied. The same mechanism works in reverse when immigration is restricted. The most striking example is the end of the Bracero Program. From 1942 to 1964, the United States admitted large numbers of Mexican agricultural workers on temporary contracts. When the programme ended at the end of 1964, supporters of the change argued explicitly that it would raise the wages and employment of domestic farm workers. Michael Clemens, Ethan Lewis and Hannah Postel tested that claim in a 2018 paper in the American Economic Review. Comparing states that had relied heavily on bracero labour with states that had not, they found no evidence that the exclusion raised wages or employment for domestic farm workers. Instead, farms in heavily exposed states mechanised faster, adopting technologies such as the mechanical tomato harvester, and shifted away from crops that could not be mechanised. The labour that had been removed was replaced by machines, not by native workers. A similar story has been documented for the immigration quotas the United States imposed in the 1920s. Research by Ran Abramitzky, Leah Boustan and colleagues found that areas more exposed to the restrictions did not see higher native earnings as a result. Agricultural areas shifted towards more capital-intensive production and mining activity contracted, while natives did not move into the vacated work. Restricting labour supply does not automatically redirect demand to natives; it can shrink or restructure the industries that employed immigrants. A worked illustration It helps to see how these adjustments combine by walking through a stylised example. The numbers here are illustrative, chosen to show the logic rather than to describe any particular city. Suppose a metropolitan area receives, over two or three years, an inflow of immigrants that increases the number of workers without a high school diploma by 10 per cent. In the simplest model, with capital, technology and the tasks each group performs held fixed, the wage of that group falls by an amount set by how responsive labour demand is to wages. If the relevant elasticity implied that a 10 per cent increase in supply lowered wages by around 3 per cent, which is roughly the order of magnitude of Borjas's national estimates, that would be the initial effect. Now allow the adjustments to operate. The newcomers spend their wages locally, raising demand for retail, housing, food and transport, and so raising demand for labour of all kinds. Native workers without diplomas, who have an advantage in English, shift towards jobs involving customer contact and coordination, where the newcomers do not compete as directly; the relevant comparison for them is no longer the whole group of workers without diplomas but the smaller group doing the same tasks. Local firms, finding labour more available, expand their premises and hire more; some delay automation projects they would otherwise have undertaken. Prices of labour-intensive local services fall a little, raising the real incomes of households that buy them. Each of these adjustments offsets part of the initial pressure. Over a few years, the wage effect for natives in the group can plausibly shrink from the initial 3 per cent to something close to zero, which is what most of the evidence reviewed in the previous chapter finds. For earlier immigrants without diplomas, who share the newcomers' language limitations and occupations, the task-shifting offset is weaker, and a larger share of the initial effect persists. For renters in a city where housing supply cannot expand, the rise in rents may be the most noticeable consequence of all. The illustration shows why estimates differ so much. A study that holds adjustments fixed, or measures effects over a very short period, will find something closer to the initial effect. A study that follows workers over a decade, or captures the changes in capital, technology and tasks, will find something closer to the net effect. Both are measuring something real; they are measuring different things. Complementarity at the top: skilled immigration and innovation The adjustments above apply chiefly to less-skilled immigration. For highly skilled immigrants the story is different and in some ways more straightforwardly positive, because skilled immigrants generate new ideas and new firms that raise productivity for everyone. Jennifer Hunt and Marjolaine Gauthier-Loiselle, writing in the American Economic Journal: Macroeconomics in 2010, found that immigrants to the United States with college degrees patented at substantially higher rates than natives, largely because they were concentrated in science and engineering. They estimated that a one percentage point increase in the share of immigrant college graduates in the population increased patents per capita by between 9 and 18 per cent. Petra Moser, Alessandra Voena and Fabian Waldinger, in a 2014 paper in the American Economic Review, studied the German Jewish chemists who fled Nazi Germany in the 1930s and found that US patenting rose sharply in the fields where these émigrés worked, driven substantially by American inventors entering those fields. The émigrés did not crowd out native scientists; they drew them in. Immigrants are also disproportionately likely to start businesses. Pierre Azoulay, Benjamin Jones, Daniel Kim and Javier Miranda, in a 2022 paper in American Economic Review: Insights, used administrative data to show that immigrants in the United States were roughly 80 per cent more likely than natives to found a firm, and that immigrant-founded firms of every size created jobs. In their framing, immigrants act more as job creators than as job takers. The pattern is visible in the list of large American technology companies with at least one immigrant founder. The effects of skilled immigration on skilled native workers are less settled. Some studies of the H-1B visa programme, which admits skilled temporary workers, find positive effects on native employment in affected firms and cities; others find evidence of displacement of natives within particular firms, or wage pressure in specific occupations such as computer programming. The balance of evidence is that skilled immigration raises aggregate innovation and productivity, but that particular groups of skilled natives can face competition. Prices, households and the labour of women Immigration also affects prices, and this is a channel through which natives gain that is rarely noticed. Immigrants are concentrated in certain services: childcare, cleaning, gardening, food preparation, construction, elder care. Patricia Cortés, in a 2008 study in the Journal of Political Economy, found that less-skilled immigration to US cities lowered the prices of immigrant-intensive services. Because these services are consumed disproportionately by higher-income households, the price effect benefits those households most. A further consequence follows. When childcare and household services become cheaper and more available, people who would otherwise provide those services at home can work more in the market. Cortés and José Tessada found in a 2011 study that less-skilled immigration increased the hours worked by highly educated women in the United States, particularly those in occupations with long-hours demands such as law and medicine. Immigration, in other words, can raise native labour supply and productivity at the top of the distribution by relieving the burden of domestic work. Similar findings have been reported for Italy and Spain. Where the costs concentrate The mechanisms above explain why average effects are small. They also explain where the costs fall, because not every group benefits from each adjustment. The first group is earlier immigrants. New arrivals are most similar to the immigrants who came before them: they share languages, occupations, networks and often neighbourhoods. The task specialisation that protects natives does little for earlier immigrants who have not yet moved into communication-intensive work. Ottaviano and Peri's estimates suggested that most of the negative wage effect of new immigration falls on previous immigrants, and this is consistent with a wide range of other studies. The irony is political: the people most exposed to new immigration are often immigrants themselves. The second group is natives with the least education and fewest options to upgrade. A worker without language advantages, without the credentials to move into supervisory work, or in a region where capital is slow to expand, may experience immigration as straightforward competition. The effects documented in the German border case and in Turkey's informal sector fell precisely on such groups. The third is workers in the informal economy and those with weak bargaining power. Undocumented immigrants typically earn less than documented workers with similar characteristics, partly because they cannot easily change employers or complain about violations of labour law. Their presence may depress wages and standards in the sectors where they work. Studies of the 1986 US legalisation programme found that legal status raised wages for those who received it, suggesting that irregular status itself suppresses wages. This is an argument less about immigration as such than about how legal status shapes bargaining power, a point that returns in the final chapter. The fourth is housing. Immigrants need somewhere to live, and in places where housing supply is constrained, more demand raises rents and prices. Albert Saiz, in a 2007 study in the Journal of Urban Economics, found that an inflow of immigrants equal to 1 per cent of a US city's population was associated with increases in average rents and house values of around 1 per cent. This benefits property owners and harms renters, including less-skilled natives and earlier immigrants. Where local planning restricts building, the effects are larger and more persistent. The housing channel is one of the most concrete ways in which immigration can impose real costs on particular residents, and it is often more politically salient than wages. The fifth is timing. Almost every adjustment described in this chapter takes time. Capital is installed gradually; natives change occupations over years; technology choices are made at the point of new investment. A sudden inflow, such as a refugee crisis, produces short-run effects that are larger than those of an equivalent steady flow. This is one reason why refugee arrivals often generate more acute local pressures than labour migration of similar size. The upshot The labour market is far more adaptive than the simple model assumes. Immigrants increase demand as well as supply; they specialise in different tasks; firms invest and choose technologies in response; prices change; natives change what they do. Taken together, these adjustments explain why the average effect of immigration on native wages is small and why, over a decade or more, it may be positive. But the adjustments are uneven. Those who cannot upgrade, earlier immigrants, informal workers, renters in constrained housing markets, and communities that experience sudden inflows can bear real costs. Those costs are not a reason to doubt the aggregate gain, which is large. They are a reason to design policies, from housing supply to language training to the rights of migrant workers, that spread the gain more widely. A labour market that absorbs newcomers well on average can still leave particular people worse off, and it is those people, not the averages, that shape the politics. Hashtags: #TheEconomicsOfImmigration #LaborMarkets #InternationalMigration #PlacePremium #MigrationSurplus #NativeWages #LaborMarketAdjustment #NaturalExperiments #MigrantSelection #TaskSpecialization #CapitalAdjustment #TechnologyChoice #HighSkilledImmigration #ImmigrantEntrepreneurship #FiscalImpact #BrainDrain #BrainGain #Remittances #MigrationAndDevelopment #EconomicAssimilation #IntergenerationalMobility #MigrantIntegration #MigrationPolicy #DistributionalEffects #FutureOfGlobalMigration
- The Economics of Philanthropy (Effective Altruism, Endowments, and Impact Investing)
Download the Book (PDF): Introduction Every year, Americans give away something close to six hundred billion dollars. Giving USA, the longest-running tally of American philanthropy, put the figure for 2024 at $592.5 billion, roughly the annual economic output of Sweden. Add the assets sitting in private foundations, university endowments and donor-advised funds, and the stock of capital earmarked for charitable purposes in the United States approaches three trillion dollars. The comparable pools in Britain, Germany, Switzerland and the Gulf states are smaller but growing, and they are organised around the same basic instruments. This is a great deal of capital. It is also capital of an unusual kind, and the unusual thing about it is the subject of this book. In most of the economy, money is allocated by a price system that is crude but relentless. A firm that turns capital into goods people do not want loses money, and eventually loses the capital. A bond issuer that cannot pay its coupons finds its borrowing costs rising until it stops borrowing. Markets misprice things constantly, but the errors generate their own correction, because someone who spots a mispricing can profit by acting against it. The feedback is built into the transaction. Charitable capital has no such feedback. The person who pays for a malaria net is not the person who sleeps under it. The donor who funds a scholarship programme does not experience its quality, and the student who benefits did not choose it from a menu of competing programmes at a posted price. When a charity does its work badly, nothing in the structure of the gift forces the donor to find out. Donors are, in the economist's language, purchasing a good they never consume on behalf of people who never pay. The result is a market in which the quantity supplied is set by the preferences of the buyers and the quality delivered is, for the most part, invisible to them. That observation is the controlling idea of what follows. Philanthropy is a system for allocating capital without prices, and almost every institution, method and controversy in the field is best understood as an attempt to supply the missing discipline, or as a fight over who gets to allocate in its absence. Effective altruism is an attempt to build a shadow price for doing good, expressed in dollars per life saved or per year of healthy life. Social return on investment is an attempt to express social value in the vocabulary of finance so that it can be compared with cost. Endowments are a set of rules about when capital should be spent, which is itself a pricing decision about the present against the future. Tax incentives are a price, set by the state, that makes the donor's gift cheaper than it looks. Impact investing is an attempt to reattach the discipline of returns to money meant to do good. And the critique of billionaire philanthropy is, at its heart, a question about who holds the power to allocate subsidised capital when no market and no electorate decides. Seeing the field this way does three useful things. First, it explains why measurement is at once indispensable and treacherous. If there is no price, something must stand in for one, and every stand-in distorts. A cost-per-life figure makes the best global health charities look astonishingly efficient and makes almost everything else invisible. A social return ratio turns a youth employment programme into a number that can be compared with a bond, and then invites the programme's managers to optimise the number rather than the youth. The question is never whether to measure but how much weight a measure can bear before it starts to bend what it measures. Second, it clarifies what the tax system is actually doing. A charitable deduction is not merely a reward for virtue. It is a public subsidy, delivered through the tax code, whose size rises with the donor's income and whose destination is chosen entirely by the donor. When a wealthy donor in the top bracket gives a dollar, the Treasury forgoes a substantial share of that dollar in tax, and the donor decides where the whole dollar goes. That arrangement may be defensible, and later chapters argue that in part it is. But it cannot be evaluated until it is seen for what it is. Third, it reframes the argument about timing. A foundation that promises to exist forever and spends five percent a year is making an implicit claim that its dollar spent in 2090 will do as much good as its dollar spent today. A donor who parks money in a donor-advised fund, takes the deduction immediately and grants it out slowly is making a similar claim, often without noticing. Whether those claims are true depends on the rate at which the cost of doing good rises or falls over time, which is an empirical question that the field has rarely asked out loud. The book proceeds in eight chapters. The first sets out the economics of giving without prices: what motivates donors, why the usual market correctives are absent, and what the scale of the charitable economy looks like. The second examines effective altruism, the most ambitious modern attempt to impose cost-effectiveness discipline on giving, including its genuine achievements and the crisis of confidence that followed the collapse of FTX in 2022. The third turns to social return on investment and the wider family of impact metrics, working through how the calculations are built and where they fail. The fourth is about endowments and private foundations: the theory of intergenerational equity, the practice of spending rules and the new tiered excise tax on the richest university endowments. The fifth dissects the tax incentives for giving and the extraordinary rise of donor-advised funds, now holding more than three hundred billion dollars in assets. The sixth assesses impact investing, from programme-related investments to social impact bonds, and the hard question of whether investing for good changes anything that would not have happened anyway. The seventh takes the critique of billionaire philanthropy seriously on its own terms and sets it against the strongest defences. The eighth draws the threads into a set of practical principles for donors, institutions and policymakers. A word on scope. The institutional detail is mainly American, because the United States has by far the largest and best documented charitable sector and because its tax rules have been copied, adapted or reacted against almost everywhere else. Where British or European practice differs in ways that matter, the text says so. The book is not a guide to choosing a charity, although readers who give will find the later chapters useful in that respect. Nor is it a history of philanthropy, though history appears where it explains why the institutions look as they do. Its concern is narrower and, I think, more important: how charitable capital is allocated, why that allocation is so hard to discipline, and what can reasonably be done about it. Readers will not find a verdict that philanthropy is good or bad. That question is too large to be useful. They will find instead an argument that the quality of giving depends on the institutions that channel it, that those institutions encode choices about measurement, timing and power, and that most of those choices can be made better than they are now. Chapter 1: A Market Without Prices Why people give Economists came late to philanthropy, and when they arrived they found it awkward. The standard model of the rational individual maximising their own welfare has trouble explaining why anyone would hand money to a stranger. The first serious attempts to fit giving into economic theory treated it as a contribution to a public good. If I care about poverty being reduced, then poverty reduction is something I value, and my gift to an anti-poverty charity is a purchase of that value. The difficulty is that poverty reduction, like clean air, benefits me whether or not I pay for it. If others are giving, I can enjoy the result without contributing. Pure public-good models therefore predicted that giving would collapse as the number of potential donors grew, and that every dollar of government spending on a cause would displace almost a full dollar of private giving to it. Neither prediction matches what we observe. Millions of people give, and government grants to charities displace only a fraction of private donations. The most influential repair came from James Andreoni in a pair of papers in 1989 and 1990. Andreoni proposed that donors are moved by two things at once: an interest in the outcome, and a private satisfaction in the act of giving itself, which he called warm glow. A donor with warm glow is an impure altruist. She cares that the hungry are fed, but she also cares that she was the one who helped feed them. Warm glow explains why giving persists in large populations, why government spending crowds out private giving only partially, and why people give to causes where their marginal contribution is negligible. It also explains something less flattering. If a large part of the reward for giving is the feeling of having given, then the donor is being paid, in emotional currency, at the moment the cheque is written, not at the moment a life is improved. The reward arrives before the result, and it arrives whether or not the result ever does. Later research has added texture without overturning the basic picture. Donors respond to being asked, and the identity of the asker matters: people give more when a friend or colleague solicits them than when a stranger does. They respond to social visibility, giving more when gifts are publicised and clustering around the thresholds at which donor lists move them into a more prestigious category. They respond to matching offers, though experimental evidence suggests that the existence of a match matters more than its generosity; a one-to-one match raises giving about as much as a three-to-one match. They respond to identifiable victims far more than to statistics, a pattern psychologists have documented repeatedly since the 1990s. A single named child in a photograph can raise more than a description of the millions like her. None of this is irrational in the sense of being incoherent. It is simply a description of preferences that are only loosely tied to impact. The donor's utility depends on the act of giving, on the esteem of peers, on a connection to the cause, on the story. It depends only weakly and indirectly on what the money actually achieves, because what the money achieves is, for most donors, something they never see. The missing feedback loop To see why this matters, consider how an ordinary market handles quality. When I buy a coat, I bear the consequences of a bad choice. If it falls apart, I am cold and poorer, and next time I buy a different brand. Firms that sell bad coats lose customers and, eventually, go out of business. The information about quality flows back to the person who paid, and the person who paid can act on it. This loop is the core of market discipline. It does not require consumers to be well informed at the start, only that they experience the product and can switch. Philanthropy breaks this loop at two points. The first break is between payer and beneficiary. The donor pays; someone else consumes. The beneficiary of a food bank experiences its quality directly but did not pay for it and usually has no alternative supplier to switch to. The donor could switch, but experiences nothing. The second break is between action and outcome. Many charitable interventions produce results that are delayed, diffuse and difficult to attribute. Did the after-school programme cause the rise in graduation rates, or would those students have graduated anyway? Did the advocacy campaign change the law, or was the law going to change regardless? Even a diligent donor who wants to know the answer often cannot find it without evidence that costs more to gather than the gift itself. The consequence is that the charitable sector faces weak selection pressure on effectiveness. A charity survives if it can raise money, and it raises money by satisfying donors. Satisfying donors is related to doing good work, but the relationship is loose. A charity that tells compelling stories, cultivates wealthy patrons and runs a polished gala can thrive while delivering little. A charity that works quietly on an unglamorous problem with enormous effect can struggle. In a commercial market, the second firm would eventually win, because its customers would notice. In the charitable market, there is no mechanism that guarantees it. This is not a new observation. Andrew Carnegie, writing in 1889, judged that of every thousand dollars spent on so-called charity, nine hundred and fifty were probably spent unwisely, and that it would be better for mankind if the millions of the rich were thrown into the sea than spent in ways that encouraged dependency. His reasoning was crude and his solution, which involved funding libraries and concert halls rather than relief, reflected his own tastes. But his diagnosis anticipated the modern one. Without some discipline on quality, giving flows to what pleases the giver. It is worth being precise about what the missing price means. It does not mean that charitable goods have no value, or that their value is unknowable. It means that there is no market-generated signal that aggregates dispersed information about value and pushes resources towards the best uses. In a functioning market, a high price for a good tells producers to make more of it and consumers to economise on it, and it does so without anyone needing to understand why the good is scarce. Charitable capital lacks that signal. Whatever information exists about which interventions work best has to be gathered deliberately, analysed centrally and communicated to donors who may or may not care. Every institution discussed in this book is, in one way or another, a device for doing that work, or for deciding who gets to do it. If the charitable market lacks price discipline, one might ask why it is organised through nonprofit institutions rather than simply through gifts from individuals to individuals, or through firms that sell charitable services on contract. Economists have offered two answers, and both bear directly on the problem of allocation. The first answer comes from Henry Hansmann, a legal scholar at Yale, whose 1980 article on the role of nonprofit enterprise remains the standard account. Hansmann called the problem contract failure. When the person paying for a service cannot observe whether it has been delivered, an ordinary profit-seeking firm has every incentive to take the money and skimp on the service. A donor who sends money to a relief agency working in a distant famine cannot check how much food reached the people it was meant for. If the agency were a firm with owners entitled to its profits, every dollar it saved by delivering less food would be a dollar in the owners' pockets. The nonprofit form addresses this by imposing what Hansmann called the nondistribution constraint: a nonprofit may earn a surplus, but it may not distribute that surplus to anyone who controls it. Managers can still be paid salaries, and they can still waste money, but they cannot pocket the difference between what donors give and what the service costs. The constraint does not guarantee quality. It removes one powerful incentive to cheat, and in doing so it makes it rational for donors to trust institutions they cannot monitor. The second answer comes from Burton Weisbrod, who emphasised the gap left by government. In a democracy, public goods tend to be supplied at the level preferred by the median voter. Citizens who want more of a particular public good than the median voter does, whether more support for a religious tradition, more funding for opera, or more research on a rare disease, cannot get it through the ballot box. The nonprofit sector allows them to supply it themselves. On this view, charities are a mechanism by which minorities with intense preferences can top up public provision without persuading the majority. Both answers are illuminating, and both contain the seed of later trouble. Hansmann's account explains why donors trust nonprofits, but it also explains why that trust is fragile: the nondistribution constraint prevents one kind of abuse while leaving the organisation free to be ineffective, self-perpetuating or captured by the preferences of its staff. It tells the donor that her money will not be stolen, not that it will be well used. Weisbrod's account explains why the sector is pluralistic, but it also frames philanthropy as a way for those who can afford it to shape the supply of public goods beyond what democratic processes would choose. When the people doing the topping up are ordinary citizens giving modest sums, that looks like healthy pluralism. When a handful of very wealthy donors can outspend the government in a policy area, as has happened in parts of American education reform and global health, it begins to look like something else. The same economic logic supports both descriptions, and much of the argument in Chapter 7 turns on where the line between them falls. The shape of the charitable economy Before turning to those devices, it helps to have the scale of the system in view. American giving is dominated by individuals, who account for about two-thirds of the total. Foundations are the second-largest source, followed by bequests and corporations. The breakdown for 2024, as estimated by Giving USA, appears in Table 1. Table 1. Sources of charitable giving in the United States, 2024. Source Amount (billions of dollars) Share of total Change from 2023 Individuals 392.45 66% +8.2% Foundations 109.81 19% +2.4% Bequests 45.84 8% -1.6% Corporations 44.40 7% +9.1% Total 592.50 100% +6.3% Source: Giving USA 2025, published by Giving USA Foundation and researched by the Indiana University Lilly Family School of Philanthropy. Shares rounded. Three features of this picture matter for the argument that follows. The first is concentration. The individual share looks democratic, but it is increasingly driven by a small number of very large donors. Research from the Lilly Family School and others has documented a long decline in the share of American households that give to charity at all, from around two-thirds at the start of the century to under half in the most recent surveys, even as total dollars have kept rising. Giving has become a larger activity carried out by fewer people. That trend is at the root of the political argument over billionaire philanthropy taken up in Chapter 7. The second is intermediation. A growing share of what counts as individual giving does not go directly to working charities. It goes to intermediaries: private foundations the donor controls, and above all donor-advised funds, which received nearly ninety billion dollars in contributions in 2024 according to the Donor Advised Fund Research Collaborative. Giving USA's own recipient data show that gifts to foundations alone amounted to nearly seventy-two billion dollars in 2024. Money given to an intermediary has been counted as charity and deducted for tax purposes, but it has not yet done anything. How long it waits, and who decides when it moves, are among the central questions of Chapters 4 and 5. The third is the distribution of destinations. Religious congregations remain the largest single recipient category, at nearly $147 billion in 2024, followed by human services and education. International affairs, the category that includes most global health and development charities, received about $36 billion. That is a small share of American giving directed at the places where, as the next chapter shows, a dollar appears to go furthest. The allocation reflects donor preferences: for local causes, for institutions with which donors have a personal connection, for religious communities that also provide social and spiritual goods to the donors themselves. Those preferences are legitimate. But they are not the result of anyone comparing the marginal value of a dollar across causes, because in a market without prices, nobody is required to. The overhead trap For decades, the most widely used stand-in for charitable quality was the overhead ratio: the share of a charity's spending that goes to administration and fundraising rather than programmes. It was easy to calculate from public tax filings, easy to understand and easy to rank. Rating agencies built their early methodologies around it, and donors absorbed the lesson that a good charity is a lean one. The overhead ratio is an instructive failure because it shows what happens when a measure is chosen for convenience rather than validity. It measures inputs, not outputs. A charity that spends ninety percent of its budget on programmes may be spending that ninety percent on programmes that do nothing. A charity that invests heavily in data systems, skilled staff and evaluation may have a high overhead ratio precisely because it is trying to find out whether it works. Worse, the pressure to keep overhead low pushes charities to underinvest in the capacities that make them effective, a pattern that researchers at the Bridgespan Group and elsewhere have called the nonprofit starvation cycle. Funders expect low overhead, charities understate their real costs to meet the expectation, funders come to believe the understated figures are realistic, and the cycle tightens. By 2013 the leaders of the three largest American charity rating organisations, GuideStar, Charity Navigator and the BBB Wise Giving Alliance, had published an open letter to donors calling the overhead ratio a poor measure of performance and urging them to look at transparency, governance, leadership and results instead. Charity Navigator subsequently rebuilt its ratings to include measures of impact and results where the data allow. The entrepreneur and activist Dan Pallotta had by then become the best-known critic of what he called the overhead myth, arguing that the sector was being prevented from growing to the scale of the problems it addressed by a moral code that treated spending on talent and marketing as waste. The lesson of the overhead episode is not that efficiency is irrelevant. It is that a measure which is easy to observe will tend to displace a measure which matters, and that an organisation evaluated on a proxy will, over time, optimise the proxy. This is Goodhart's law applied to charity, and it recurs in every attempt to supply a price where the market does not. Cost-effectiveness estimates, social return ratios, impact scores and payout rates are all, in principle, better measures than overhead. Each of them is also vulnerable to the same dynamic. The measure becomes the target, the organisation learns to hit the target, and the relationship between the target and the good it was meant to track weakens. There is a deeper point beneath this. In a market with prices, the measure of success is not chosen by anyone. It emerges from the choices of many buyers and sellers, and it is hard to game because it is the aggregate of so many independent decisions. In philanthropy, the measure has to be chosen, and whoever chooses it holds a kind of power. The rating agency that decides overhead matters shapes which charities grow. The evaluator that decides deaths averted is the right unit shapes which causes receive money. The donor who decides that her alma mater is the most deserving institution in the world shapes, at the margin, where tens of millions of dollars go. Measurement, in this sense, is never purely technical. It is a way of allocating authority over capital that no market is allocating. This is why the economics of philanthropy cannot be separated from its politics. The absence of prices creates a vacuum, and the vacuum is filled by some combination of expert judgement, donor preference and public rule. The chapters that follow examine each of these in turn: the experts who have tried to build a science of giving well, the institutions that decide how fast charitable capital moves, the state that subsidises the whole system, the investors who have tried to reattach financial returns to social ends, and the very wealthy individuals whose preferences now shape a growing share of the result. The question running through all of them is the same. When there is no price to tell us where charitable capital should go, who decides, by what standard, and with what accountability? The natural starting point is the group that has answered that question most confidently: the effective altruists, who argued that the missing price could, in large part, be calculated. Chapter 2: Pricing the Good: Effective Altruism and Cost-Effectiveness From a drowning child to a research agenda In 1972 the philosopher Peter Singer published an essay called "Famine, Affluence, and Morality", written against the background of the humanitarian catastrophe in what was then East Bengal. Its central argument rested on a simple analogy. If you walk past a shallow pond and see a child drowning, you ought to wade in and save her, even if it ruins your clothes. The cost to you is trivial compared with the loss of a life. Singer then asked what distinguished the drowning child from a child dying of preventable causes on the other side of the world. Distance, he argued, is not morally relevant. Nor is the fact that many others could help too. If we can prevent something very bad from happening without sacrificing anything of comparable moral importance, we ought to do it. The implication was that affluent people in rich countries were obliged to give far more than they did, and that the ordinary distinction between duty and charity was indefensible. For three decades the essay was mainly an argument in moral philosophy seminars. What turned it into a movement was the addition of an empirical claim: that the difference in effectiveness between charities is not modest but enormous, often a factor of ten or a hundred, and that it can be estimated. If that is true, then the choice of where to give matters more than the choice of how much. A donor who gives a modest sum to the most effective charity may do more good than one who gives ten times as much to an ordinary one. The institutional form arrived in the late 2000s. In 2007 two former hedge fund analysts, Holden Karnofsky and Elie Hassenfeld, founded GiveWell after discovering that the charities they approached could not tell them, in any rigorous way, what their donations would accomplish. GiveWell set out to answer that question itself, researching a small number of charities in depth and publishing its reasoning in full. In 2009 the Oxford philosophers Toby Ord and William MacAskill founded Giving What We Can, whose members pledge to give at least ten percent of their income to effective charities. The phrase "effective altruism" was adopted around 2011 as an umbrella for these and related efforts, and by the time MacAskill's "Doing Good Better" and Singer's "The Most Good You Can Do" appeared in 2015, the movement had its own vocabulary, conferences and, increasingly, its own money. The largest source of that money was Good Ventures, the foundation of the Facebook co-founder Dustin Moskovitz and Cari Tuna, which partnered with GiveWell to create what became Open Philanthropy, a grantmaker that renamed itself Coefficient Giving in November 2025 and reports having directed more than four billion dollars in grants across its history. In the terms of this book, effective altruism was the most serious attempt yet to calculate the missing price. Its premise was that the value of a charitable intervention could be expressed in a common unit, that interventions could be ranked by the cost of producing a unit of that value, and that donors should buy the cheapest units first. The arithmetic of cost-effectiveness GiveWell's method is worth describing in some detail, because it is the most transparent version of the approach and because its strengths and weaknesses are visible in the calculations themselves. Take an insecticide-treated bednet campaign of the kind run by the Against Malaria Foundation. The analysis begins with cost: what it costs to buy, ship and distribute a net, including the costs borne by governments and other partners, not just the charity's own spending. It then asks how many nets reach people, how many are used, how long they last and how many people sleep under each. From randomised trials of bednets, most of them conducted in the 1990s and summarised in systematic reviews, it takes an estimate of how much a net reduces child mortality from malaria. That effect is adjusted for present-day conditions: malaria burden in the specific region, the spread of insecticide resistance among mosquitoes, the fact that other interventions such as seasonal preventive medication may already be reducing deaths. The result is an estimate of deaths averted per dollar. GiveWell adds estimates of other benefits, such as the long-term income gains that follow from reducing childhood illness, and expresses the whole in a common unit of value so that a bednet programme can be compared with, say, a programme that pays families small cash incentives to vaccinate their children. The headline results are striking. GiveWell's published estimate is that it typically costs between $3,000 and $5,500 to save a life through its top charities, with the cheapest opportunities in high-burden areas at the lower end of that range. GiveWell itself warns that the figure varies widely by location, that it expects the cost to rise over time as the cheapest opportunities are funded, and that the headline number is a simplification of a much more detailed model. Even allowing for those caveats, the gap between these figures and the implied cost per life of typical public health spending in rich countries, often measured in millions of dollars per statistical life, is large enough to explain why the argument has been persuasive to so many. The scale has grown accordingly. In its 2024 metrics year GiveWell directed about $397 million to the programmes it recommends, from more than 30,000 donors. It estimates that the grants made that year will save around 74,000 lives, most of them children in sub-Saharan Africa. The largest single recipient, the Against Malaria Foundation, received about $150 million. Two features of this method deserve emphasis. The first is that it depends on evidence of a particular kind. The effectiveness of bednets, deworming pills, vitamin A supplementation and seasonal malaria chemoprevention is known because these interventions have been tested in randomised controlled trials, the method whose use in development economics earned Abhijit Banerjee, Esther Duflo and Michael Kremer the Nobel Memorial Prize in 2019. The approach privileges interventions that can be tested this way, which tend to be discrete, deliverable and measurable over a few years. The second feature is that the method is explicitly marginal. GiveWell does not ask whether bednets are good in general. It asks what the next dollar will buy, given what other funders are already paying for. When a charity has more money than it can use well, GiveWell stops recommending donations to it, even if its programme remains excellent. The comparison is anchored to a benchmark that GiveWell chose with care: unconditional cash transfers to very poor households, of the kind delivered by the charity GiveDirectly. The logic is appealing. Cash is the simplest thing a donor can give, it respects the recipient's own judgement about what she needs, and its effects have been studied in randomised trials in Kenya and elsewhere since the early 2010s, including a well-known evaluation by Johannes Haushofer and Jeremy Shapiro. If a more complicated intervention cannot beat cash, it is hard to justify the paternalism of choosing on the recipient's behalf. GiveWell accordingly expresses the cost-effectiveness of its recommended programmes as multiples of its estimate for cash transfers, and it has for years funded only opportunities that clear a bar set well above that benchmark. Cash functions, in effect, as the index fund of global giving: the default against which any active choice must justify itself. The benchmark has itself moved. In late 2024 GiveWell published a re-evaluation of the evidence on cash transfers, drawing on newer studies that found larger spillover benefits to neighbouring households and local economies than earlier work had captured. It concluded that GiveDirectly's programme was three to four times more cost-effective than it had previously estimated, while judging that its top charities remained at least twice as cost-effective as the revised figure. The episode is a small illustration of a larger point. A shadow price built from research is only as stable as the research, and when the evidence improves, the ranking of everything measured against it can shift. The events of 2025 tested the marginal logic in a different way. When the United States government abruptly cut large parts of its foreign aid programme early that year, many of the malaria, tuberculosis and nutrition programmes that effective altruist donors had treated as someone else's responsibility suddenly faced funding gaps. The marginal dollar, which had previously bought whatever was left after governments and large multilateral funders had paid for the core, was now being asked to replace some of that core. GiveWell reported approving 131 grants worth $418 million in 2025, of which about $53 million was aimed specifically at urgent needs created by the aid cuts. The response was quick and careful, but its scale also exposed a hard truth about philanthropic capital. Even the best-organised private funders in global health command sums that are small relative to what governments spend, and the cost-effectiveness of a charitable dollar depends heavily on decisions made by public budgets that philanthropy does not control. That marginal logic is exactly what a price system provides in a market. A price tells a buyer what the next unit costs, not what the average unit was worth. GiveWell's analysis is, in effect, an attempt to compute the marginal cost of a unit of good and to publish it where donors can see it. What the calculations miss The criticisms of this approach fall into three groups, and they are worth separating because they have very different implications. The first group concerns the reliability of the numbers. Cost-effectiveness estimates are built on chains of assumptions, each uncertain, and the uncertainty compounds. A modest error in net usage, net durability, baseline mortality and the transferability of trial results can shift the final figure by a large factor. The best-known illustration is deworming. A study by Edward Miguel and Michael Kremer, published in 2004, found that school-based deworming in western Kenya reduced absenteeism, and long-run follow-ups suggested large gains in adult earnings. In 2015 a reanalysis of the original data by epidemiologists at the London School of Hygiene and Tropical Medicine found coding errors and argued that the education effects were weaker than claimed, while a Cochrane systematic review found little consistent evidence that mass deworming improved nutrition or school performance. The exchange became known as the worm wars. GiveWell continued to recommend deworming, but explicitly on the basis that the possible long-run benefit was large enough to justify a bet even though the evidence was uncertain. That is a reasonable judgement, but it is a judgement, not a measurement. Once the most celebrated numbers in effective altruism are recognised as expected values under deep uncertainty, the distance between calculated price and informed opinion narrows. The second group concerns what the numbers leave out. The cost-effectiveness framework is strongest where outcomes are countable and weakest where they are not. It handles child deaths from malaria well. It handles the strengthening of a national health system, the growth of a civil society organisation, or a change in a country's governance badly, because those outcomes are hard to attribute and hard to value. The economist Angus Deaton has argued that aid which bypasses governments may weaken the accountability relationship between states and their citizens, and that this cost does not appear in any charity's spreadsheet. Leif Wenar, a philosopher, pressed a related point in a widely read 2024 essay, arguing that effective altruism's calculations systematically ignore the harms and side effects of interventions and give donors a false sense of precision. One need not accept every part of these critiques to see the structural problem. A shadow price computed only for goods that can be measured will pull money towards those goods and away from everything else, regardless of whether the unmeasured goods are more valuable. This is the overhead problem in a more sophisticated form. The measure is far better than the overhead ratio, because it tracks outcomes rather than inputs. But it shares the vulnerability of any partial price. It will allocate efficiently among the things it can see and neglect the things it cannot. The third group of criticisms is philosophical, and concerns the choice of unit. To compare a bednet with a cash transfer, one has to decide how much an averted death is worth relative to a given increase in consumption, and how the value of saving a child's life compares with the value of saving an adult's. GiveWell publishes its moral weights and has at various points surveyed beneficiaries in low-income countries about how they would make these trade-offs, which is an honest attempt to ground the numbers in something other than the analysts' intuitions. But no survey resolves the underlying question. The common unit is a construction, and different reasonable constructions produce different rankings. None of these objections shows that cost-effectiveness analysis is useless. On the contrary, the evidence that some interventions are many times more cost-effective than others is robust across a wide range of assumptions. What the objections show is that the calculation cannot bear the full weight that the early rhetoric of effective altruism placed on it. It is a tool for narrowing choices within a domain where outcomes are measurable. It is not a universal price. Longtermism, FTX and the limits of expected value The movement's most consequential intellectual turn came in the late 2010s, when a growing number of its leaders argued that the most important thing philanthropy could do was reduce the risk of catastrophes that might end or permanently cripple human civilisation: engineered pandemics, nuclear war and, above all, misaligned artificial intelligence. The reasoning was an extension of the same expected-value arithmetic. If the future could contain vastly more people than the present, then even a tiny reduction in the probability of extinction would, in expectation, be worth more than any quantity of present-day bednets. MacAskill's "What We Owe the Future", published in 2022, made the case for this longtermist view to a general audience. The trouble with this move, from the standpoint of allocation, is that it detached the calculation from evidence entirely. A bednet estimate can be checked against trials, mortality data and distribution surveys. An estimate of how much a research grant reduces the probability of human extinction cannot be checked against anything. The expected value depends almost wholly on the probabilities one assigns, and those probabilities are, in most cases, guesses. The method that had been developed to discipline donors' preferences was now being used to license preferences that were, in practical terms, undisciplined. Critics inside and outside the movement noted that the causes it favoured, particularly those related to artificial intelligence, coincided closely with the interests and anxieties of the technology industry from which much of its money came. The collapse of the cryptocurrency exchange FTX in November 2022 turned these criticisms into a crisis. Its founder, Sam Bankman-Fried, had been the movement's most prominent new donor, having embraced the idea of "earning to give", making as much money as possible in order to give it away. His FTX Future Fund had committed large sums to longtermist causes. When the exchange failed amid revelations that customer deposits had been diverted, the Future Fund's team resigned, grants were left unpaid or subject to recovery by bankruptcy administrators, and Bankman-Fried was later convicted of fraud and, in March 2024, sentenced to twenty-five years in prison. The episode did not show that the ideas of effective altruism were wrong. It did show that a movement which had become heavily dependent on a small number of very large donors was exposed to their failures, and that the confident use of expected-value reasoning could coexist with a remarkable lack of scrutiny of where the money came from. The aftermath has been instructive. Much of the effective altruist ecosystem has since reaffirmed its commitment to evidence-backed global health work, where the original methods are strongest. Coefficient Giving now runs pooled funds that other donors can join, covering areas from global health to animal welfare to the safety of artificial intelligence, and it has sought to diversify beyond its founding donors. GiveWell's own donor base is broader than it was a decade ago. The movement is, in other words, correcting for the concentration problem that any philanthropic institution faces when a single funder's preferences dominate. The lasting contribution of effective altruism to the economics of philanthropy is best stated modestly. It demonstrated that charitable options differ enormously in cost-effectiveness, that careful analysis can identify some of the best, and that donors will respond to that analysis if it is published transparently. It built, for one corner of the charitable market, something close to a functioning price signal. Its failures came when it treated that signal as a universal measure of value, when it extended expected-value reasoning into domains where the probabilities are unknowable, and when it allowed a handful of donors to define the agenda. Those are not failures peculiar to effective altruism. They are the characteristic hazards of every attempt to price the good, and the next chapter examines them in a very different setting: the attempt, popular among governments and foundations, to express social value as a financial return. Chapter 3: Social Return on Investment and the Limits of Measurement Social value in the language of finance Effective altruism tried to build a price for the good by estimating outcomes in a single unit of welfare. A parallel movement, with different origins and a different constituency, tried to do something that sounds similar but is subtly different: to express social value in money, so that a charitable programme could be evaluated with the same tools used for a financial investment. Its best-known method is social return on investment, or SROI. SROI grew out of work in the late 1990s at the Roberts Enterprise Development Fund, now known simply as REDF, a San Francisco philanthropy founded by the investor George Roberts to fund social enterprises that employ people facing barriers to work, such as homelessness, addiction or a criminal record. REDF wanted to know whether its grants were paying off, and it borrowed its vocabulary from venture capital. If a social enterprise moved a formerly homeless man into stable employment, what was that worth? REDF's answer was to add up the monetary consequences: the reduction in welfare payments and public services he no longer used, the taxes he now paid, the income he earned. Those figures could be compared with the grant, producing a ratio of social return to investment. The method was developed further in Britain in the 2000s, where it found a sympathetic audience among governments interested in commissioning public services from charities and social enterprises and wanting a common language for comparing them. The Cabinet Office's Office of the Third Sector published "A Guide to Social Return on Investment" in 2009, later revised, and the approach is now stewarded by Social Value International and its national affiliates, which set out a list of principles and accredit practitioners. The Public Services (Social Value) Act of 2012 required public bodies in England and Wales to consider wider social, economic and environmental value when commissioning services, which gave the language of social value, if not SROI specifically, a place in procurement. The appeal of SROI is obvious. A funder faced with a youth mentoring programme, a community garden and a debt advice service cannot easily compare them. If each can produce a ratio, the funder can rank them. And a ratio of, say, four to one, meaning four pounds of social value created for every pound invested, is a sentence that a board of trustees or a government minister can grasp. How an SROI is built An SROI analysis proceeds through a sequence of steps that is simple to describe and hard to do well. It begins by identifying the stakeholders affected by an activity and mapping how the activity changes their lives, a chain running from inputs to outputs to outcomes. It then asks how each outcome can be measured and how it can be given a monetary value. Some outcomes have obvious monetary values, such as higher wages or lower spending on benefits. Others require proxies. The value of improved wellbeing might be estimated from the amount people are willing to pay for comparable improvements, or from wellbeing valuation techniques that use survey data to infer how much income would produce an equivalent rise in life satisfaction. The crucial steps come next, because they address the question of what the activity actually caused. An SROI analysis applies four adjustments. Deadweight is the share of the outcome that would have happened anyway. If forty percent of participants in an employment programme would have found jobs without it, only sixty percent of the observed employment can be credited. Attribution is the share of the outcome caused by others. If participants were also receiving help from a job centre and a family member, some of the credit belongs there. Displacement is the share of the outcome that simply moved a problem elsewhere. If a programme helped its participants into jobs that would otherwise have gone to other unemployed people, the net gain to society is smaller than the gain to participants. Drop-off is the rate at which the outcome fades over time. The value in year one may be sustained, but by year three much of it may have gone. Once these adjustments are applied, the remaining value in each future year is discounted back to the present, usually using the government's recommended social discount rate, and divided by the investment. The result is the SROI ratio. A stylised worked example shows how much turns on the adjustments. Suppose a charity spends £200,000 on an employment programme for fifty young people who have been out of work for more than a year. Thirty of them find jobs. Suppose also that a year of employment for each of them is valued at £12,000, combining higher earnings, lower benefit payments and a monetised estimate of improved wellbeing. The gross first-year value is £360,000, which already looks like a return of 1.8 to one, and if the benefit is assumed to last for three years without discounting, the headline ratio reaches 5.4 to one. Now apply the adjustments. If comparable young people find work at a rate that implies a third of the thirty would have found jobs anyway, deadweight removes a third. If a quarter of the remaining credit belongs to other services, attribution removes a further quarter. If a fifth of the jobs displaced other job-seekers, displacement takes a fifth of what remains. And if the effect drops off by a third each year, the three-year stream shrinks. Working through these figures, the first-year value falls from £360,000 to £144,000, and the three-year total, even before discounting, comes to about £304,000: £144,000 in the first year, £96,000 in the second and £64,000 in the third. The ratio falls from 5.4 to about 1.5. Every one of those adjustment rates is a judgement that could reasonably be halved or doubled. A different analyst, with defensible choices, could report a ratio of one or a ratio of four from exactly the same programme. It is instructive to set this against the most celebrated monetised social return in the research literature. The Perry Preschool Project, run in Ypsilanti, Michigan in the 1960s, randomly assigned a small group of disadvantaged three- and four-year-olds to a high-quality preschool programme and followed them, along with a control group, for decades. Because assignment was random, the control group supplied a genuine counterfactual: deadweight could be measured rather than assumed. The economist James Heckman and colleagues, in an analysis published in 2010, estimated the programme's social rate of return at roughly seven to ten percent a year, driven largely by reduced crime and higher earnings, after correcting for problems in earlier estimates. That is a respectable return rather than a spectacular one, and it is far more credible than most SROI ratios precisely because it rests on a controlled comparison followed for forty years. Even so, it has been debated vigorously, partly because the sample was small, partly because the valuation of avoided crime drives so much of the result, and partly because a programme run in one town in the 1960s may not tell us much about programmes run elsewhere today. If the best-evidenced social return in the literature attracts that much argument, a ratio produced from a single year of programme data and assumed adjustment rates deserves considerably more scepticism. The discount rate matters too. A benefit that lasts many years is worth less today than its undiscounted sum, and the choice of rate can change the ratio materially for long-lived outcomes. In Britain, the Treasury's Green Book recommends a social discount rate of 3.5 percent for most public appraisal, and SROI practitioners commonly follow it. A lower rate flatters programmes whose benefits are distant, such as early childhood interventions; a higher one flatters programmes whose benefits are immediate. The rate is a statement about how much the future matters, and like the other parameters it is chosen rather than discovered. The same question, in a much larger form, sits at the heart of the next chapter's discussion of perpetual endowments. The worked example above is not an artificial case of carelessness. It illustrates the central fact about SROI: the ratio is highly sensitive to parameters that are rarely observed directly. Deadweight in particular requires a counterfactual, an estimate of what would have happened without the programme, and that estimate is exactly what most charities lack. A randomised trial can supply it, but few programmes have one. In its absence, deadweight is often taken from national statistics, from the literature on comparable programmes, or from the analyst's judgement. What SROI measures, and what it rewards The honest practitioners of SROI have always acknowledged these limitations. Social Value International's principles stress involving stakeholders, not over-claiming and being transparent about assumptions, and a well-conducted SROI report sets out its adjustments explicitly and tests how sensitive the result is to them. At its best, the process is more valuable than the ratio. Mapping how a programme is supposed to change lives, asking which of those changes can be observed and asking what would have happened otherwise are exactly the questions a charity ought to ask about itself. The problem arises when the ratio travels without the report. A headline figure of "£5 of social value for every £1 invested" is quoted in annual reviews, funding applications and press releases, stripped of its assumptions. Reviews of published SROI studies have repeatedly found wide variation in method and a tendency towards high ratios, which is what one would expect when the organisation commissioning the study is also the organisation being evaluated. Because there is no standard counterfactual, no standard set of proxies and no independent audit of most studies, ratios from different organisations are rarely comparable. The very thing that made SROI attractive, a single number that allows comparison across causes, is the thing it is least able to deliver. There is also a subtler distortion. SROI tends to value outcomes by their fiscal consequences, because those are the easiest to monetise. A programme that keeps someone out of prison generates a large fiscal saving. A programme that helps an elderly person feel less lonely generates a small one, unless the analyst uses wellbeing valuation, which introduces a separate layer of assumptions. The method therefore tends to favour interventions whose benefits accrue to the state, and to favour populations whose problems are expensive to the state. That is a coherent perspective for a government commissioner deciding which services to buy. It is not obviously the right perspective for a philanthropist trying to do the most good, and the two are easily confused. The comparison with effective altruism is illuminating. Both approaches try to supply a price where none exists. The effective altruist approach, as practised by GiveWell, typically compares interventions within a narrow domain, relies heavily on trial evidence, and expresses results in units of welfare such as deaths averted. SROI is used across a wide range of domains, relies on whatever evidence is available and expresses results in money. The first is more rigorous and narrower; the second is broader and more pliable. Neither escapes the problem that a price calculated by an interested party is not the same as a price discovered by a market. The wider family of impact measures SROI is one of several approaches that philanthropists and social investors use to stand in for the missing price, and it helps to see them side by side. The main approaches differ in their unit of value, the evidence they require, the domains they suit and the characteristic way they go wrong, as Table 2 sets out. Table 2. Four approaches to measuring charitable impact. Approach Unit of value Evidence typically used Best suited to Characteristic failure Cost-effectiveness analysis Deaths averted, DALYs or welfare units per dollar Randomised trials, epidemiological data Global health, discrete interventions Neglects what cannot be measured Social return on investment Money value per unit of money invested Programme data, proxies, stakeholder input Local services, social enterprise Ratios inflated by favourable assumptions Impact metrics and ratings (such as IRIS+ or B Impact Assessment) Standardised indicators and scores Self-reported organisational data Impact investment portfolios, enterprises Outputs reported instead of outcomes Trust-based or participatory grantmaking No common unit; judgement of grantees and communities Relationships, qualitative learning Advocacy, systems change, community work Weak accountability for results The first three rows are recognisably attempts to construct prices. The fourth is a deliberate decision not to. Trust-based philanthropy, which has grown rapidly in the United States over the past decade, holds that the reporting burdens imposed by funders are expensive, distorting and often useless, and that grantees and the communities they serve are better placed to judge what works than distant funders with spreadsheets. It favours unrestricted, multi-year grants with light reporting. The most prominent practitioner is MacKenzie Scott, who has given away tens of billions of dollars since 2019, largely in the form of large unrestricted gifts to organisations identified through quiet research, with no application process and minimal reporting requirements. Trust-based grantmaking can look like a rejection of the whole project of measurement, but it is better understood as a different answer to the same problem. If there is no price, someone has to decide. Cost-effectiveness analysis gives the decision to expert analysts. SROI gives it to evaluators, often hired by the organisation being evaluated. Trust-based grantmaking gives it, after the initial selection, to the grantees themselves. Each has costs. Expert analysis neglects what cannot be counted. Self-evaluation flatters. Trust without verification can let weak organisations persist. The real question for any funder is not which method is correct but which errors it is prepared to live with in a given domain. A useful discipline, whatever the method, is proportionality. The cost of measuring should bear some relation to the size of the decision it informs. A foundation deciding whether to commit fifty million dollars to scaling a programme across a country should demand serious evidence, ideally including a counterfactual. A donor giving a few hundred dollars to a local food bank does not need an SROI report, and the food bank should not be asked to produce one. Much of the frustration with impact measurement in the charitable sector comes from funders who apply the evidentiary standards of the first case to the second, loading small organisations with reporting requirements that consume a meaningful share of the grant and produce numbers nobody uses. A second discipline is to treat any single ratio with suspicion and any change in a ratio with more interest. A programme's SROI in isolation says little. The same programme's SROI measured consistently over time, or two programmes evaluated by the same independent analyst with the same assumptions, says a good deal more. Measurement works best as a tool of learning within an organisation or a portfolio, and worst as a currency for competition between organisations that each choose their own exchange rate. The broader lesson of this chapter reinforces the last one. Every attempt to supply a price for charitable value involves choices that the price itself conceals: the choice of unit, of counterfactual, of whose benefits count and of how the future is weighted. Those choices are unavoidable. They are also where the power lies. A price discovered by a market distributes that power across millions of participants. A price calculated by a funder or evaluator concentrates it. That concentration becomes even more consequential when the capital in question is designed to last forever, which is the subject of the next chapter. Hashtags: #TheEconomicsOfPhilanthropy #CharitableCapital #MarketWithoutPrices #WarmGlowGiving #ContractFailure #NondistributionConstraint #PublicGoods #EffectiveAltruism #CostEffectivenessAnalysis #ShadowPricing #GiveWell #MarginalImpact #SocialReturnOnInvestment #ImpactMeasurement #GoodhartsLaw #Endowments #IntergenerationalEquity #CharitableTaxIncentives #DonorAdvisedFunds #ImpactInvesting #SocialImpactBonds #BillionairePhilanthropy #PhilanthropicPower #PhilanthropicAccountability #FutureOfPhilanthropicCapital
Latest Book Releases:










































