What if AI systems are the forever chemicals of our time?
It was a new technology poised to change the world. Only a few scientists knew it might be slowly poisoning people.
Sometime around 1961, toxicologists at DuPont realized low doses of synthetic chemicals known as PFAS seemed to be leading to liver enlargement in rats, and subsequently found evidence of links to birth defects and prostate cancer. But the company suppressed the evidence. As a result, these toxic chemicals made their way into everything from our raincoats to our food packaging, and eventually, our water supplies and our bloodstreams. You’ve probably heard of PFAS by their more conventional name: forever chemicals.
It took decades for public scientific research to catch up. Once it did, harms associated with PFAS have come into focus: cancer, impaired vaccine response, negative impacts to human reproduction, and more.1 But data about these hazards emerged only after the chemicals had been broadly diffused.
I think something similar is happening with AI. Others have alluded to it: Jonathan Zittrain has compared AI to asbestos: extremely useful, embedded invisibly everywhere, and very hard to clean up if we discover a problem. Documentarian Arthur Jones recently pointed to herbicide Roundup, which seeped into the environment, causing significant disruptions to ecosystems as well as troubling health effects. PFAS, asbestos, Roundup — all tell the story of substances with significant industrial value that were widely adopted before their harms were fully understood, and of how their economic utility (plus a hefty dose of corporate lobbying) continues to undermine efforts to protect people from harmful levels of exposure.
As with toxic substances, while some types of AI harms may be immediately observable and highly acute, others will surely accumulate at a low dose, wreaking havoc down the line in ways we can neither imagine nor realistically measure today. Like chemicals, as well as ionizing radiation, exposure to certain types of AI may be known (or reasonably anticipated) to increase risk of harm, but not everyone exposed will experience the same outcomes or at the same likelihoods. And those who become harmed are nearly certain to be unidentifiable at an individual level.
Every analogy for AI is also an argument about who should be in charge of it.
When people compare AI to nuclear weapons, they are mostly arguing for arms control and a small priesthood of cleared experts. When they compare it to electricity, they are suggesting that innovation ought not be strangled and that infrastructure investments should be prioritized. When they compare it to social media, they are making a case for content rules and for accountability of centralized actors. Choosing the wrong analogy for AI, many have pointed out, will set us on a path that will prove to be poorly suited for the challenges ahead, and I do think all of these analogies have something to offer. But they strike me as fundamentally incomplete.
When AI policy conversations invoke chemical and radiological risk, they mostly mean the propensity of advanced AI models to help develop weapons, but the measurement science and regulatory structures that emerged around both hazards — commonly referred to as environmental health regulation — is an analogy that demands more attention. Experts are uncovering emerging evidence that the AI technologies that are taking off might be causing harm. The science of measuring that harm is immature. The harms are both acute and cumulative. And measurement is and will remain political and contested. Environmental health regulation has been wrestling with exactly this combination for fifty years.
This is the story of how U.S. scientific and regulatory communities confronted these questions in the 1970s and 80s, and how we can learn from that history as we confront remarkably similar questions about governing AI today.
A Crisis of Legitimacy
By the late 1960s, the federal government had been regulating chemicals and ionizing radiation for decades, but mostly focused on acute poisoning, the kinds of problems that show up in days or weeks. That wasn’t necessarily the wrong focus; it was triggered by evidence of obvious harms. But two things began to change at the same time, in a way that should feel familiar to anyone watching AI policy now. Detection methods were dramatically improving, finding substances at parts-per-million and then parts-per-billion levels that were previously undetectable. And the science of low-dose harm (cancer in particular but also reproductive and developmental effects) was starting to suggest that exposures once considered safe might not be. Regulators were suddenly being asked to evaluate hundreds of new chemicals, but without the data or consensus methods to do so. The resemblance to where AI policy now finds itself is not subtle.
The result was a decade of methodological discord. The 1958 Delaney Clause imposed a zero tolerance policy for food additives, forbidding approval of any that were shown to cause cancer in humans or animals, at any dose. When the FDA invoked it to ban artificial sweeteners based on rat studies linking the compounds to cancer, critics argued that rat doses and effects were unrealistic and that the law’s indifference to dosage was driving overreaction. Public health advocates countered that new measurement methods were surfacing harms that had always been there. The same fight was playing out simultaneously over asbestos, formaldehyde, and a long list of other substances, with the EPA, FDA, OSHA, the Consumer Product Safety Commission producing wildly different risk estimates from the same underlying studies because they were using different extrapolation methods, different statistical assumptions, and different default models. Industry plaintiffs sued. Public interest groups sued. Scientific advisory committees disagreed openly. By the early 1980s, the perceived legitimacy of federal risk regulation was in serious trouble.
At the direction of Congress and the FDA in response to this legitimacy crisis, the National Academy of Sciences conducted a study on the matter, resulting in a report colloquially known as the Red Book.2 The report aimed to give the federal government a coherent framework for making public health decisions in the face of uncertainty. It recognized that human health risk assessment by regulatory agencies typically requires them to address fundamental gaps in knowledge about dose-response relationships, and that do so so, they must make explicit assumptions about that relationship to facilitate decisionmaking. The committee drew a hard line between risk assessment (what is the harm and how confident are we?) and risk management (given what we know and don’t know, what should we do about it?). It made the case that collapsing the two leads to bad outcomes in both directions: assessments get tailored to serve policy goals, and policy decisions enjoy false cover of scientific certainty.
The same conceptual collapse threatens AI debates. The validity crisis of evaluations means that today, when developers argue that a model is fit for release, it’s often not clear the extent to which the conclusion was predestined by the choice of evaluation method, and the opacity of evaluation methods in general risks obscuring political or commercial calls under the guise of scientific choices.
The Red Book didn’t resolve underlying scientific questions, but it did set the terms on which the science would be argued about for the next four decades. Its four-stage framework — hazard identification, dose-response, exposure assessment, risk characterization — became the lingua franca of federal risk regulation, and shaped how agencies organized themselves internally, separating scientific staff tasked with conducting risk analysis from policy staff making “acceptable risk” decisions based on that analysis.
Naming the stages of risk assessment made it possible to disagree about them more clearly (the sort of intervention I value), but this was not the same thing as making the disagreement go away. The question of dose-response, or how to extrapolate from high-dose animal studies to the much lower exposures humans may actually encounter, remains a particularly enduring source of dispute. Should you assume a linear relationship all the way down to zero? A threshold below which exposure is essentially harmless? Something in between? Each set of assumptions embed fundamentally different default judgments about how much caution is warranted, how much exposure-related risk is acceptable, in the face of uncertainty at low doses. And each choice can produce risk estimates that differ by orders of magnitude.
Common dose-response curves
Take, for instance, the question of risk from low-dose ionizing radiation exposure. Since the 1950s, U.S. and international radiation protection has been built on the linear no-threshold (LNT) dose-response model: the assumption that any dose of ionizing radiation, however small, carries some proportional increase in the risk of cancer. The model has been reviewed and re-affirmed repeatedly by the U.S. National Academies, international scientific bodies, and the Nuclear Regulatory Commission (NRC) itself, most recently in a 2021 rejection of formal petitions to abandon it based on a counterargument holding that low doses of radiation are not just harmless but potentially beneficial. The choice of low-dose risk model has significant implications for nuclear power, medical imaging, and occupational safety, so it’s unlikely to be surprising that industry has had reason to leverage scientific uncertainty to advocate for a shift to the use of threshold models.
For a long time, these methodological arguments didn’t change radiation exposure policy. But that changed last May when President Trump signed Executive Order 14300, “Ordering the Reform of the Nuclear Regulatory Commission”, one of four executive orders issued the same day aimed at quadrupling U.S. nuclear capacity by 2050 (to power, among other things, the data centers that train and serve large AI models). Sections 1 and 5(b) of the order direct the NRC to reconsider its reliance on the LNT model and its associated as low as reasonably achievable (ALARA) approach to managing radiation exposures, declaring that these models “lack sound scientific basis and produce irrational results,” and instructing the agency to adopt “determinate radiation limits” — that is, threshold-based limits — within eighteen months, with no cited supporting evidence. The order was not responsive to new epidemiological data, and contradicted plenty of existing evidence. Rather, it directed the regulator to change its default methodological assumption, laying the groundwork to clear away barriers to desired policy outcomes. The Bulletin of the Atomic Scientists explained that the order asked the NRC to overturn “the bedrock of radiation exposure risk analysis for decades.”
A similar erosion is happening at the EPA, where the deputy administrator recently issued an internal memo instructing the agency to stop using its Integrated Risk Information System (IRIS), the program that for nearly forty years has produced the federal government’s authoritative assessments of how dangerous individual chemicals are. Going forward, the memo said, the agency’s policy offices would conduct those assessments themselves. A parallel proposed rule would weaken the methodology underpinning chemical regulation under the Toxic Substances Control Act, including the asbestos and formaldehyde rules the previous administration had tried to strengthen. As ProPublica reported, the change opens the door to weakening hundreds of regulatory protections built on IRIS findings by suggesting that the underlying assessments can no longer be trusted. Dr. Chris Frey, who led the EPA’s research arm under the Biden administration, offered a skeptical take to the reporters:
It’s a science-based agency where the science informs decisions, not that the decisions dictate the science, which is basically what they’re proposing.
These are precisely the dynamics the environmental health regulatory tradition has been wrestling with for decades: whoever controls the default assumptions controls the conclusions, without ever needing to make a political claim.
Hazards and Outrage
Environmental health regulation also offers a warning about what happens when risk experts get overly fixated on their own understandings of relevant risks, and dismiss public concern.
It’s a reasonably consensus position that AI is in need of more mature measurement science, and progress on that front is indeed critical. Valid and accurate measurement instruments are necessary to inform what AI risks matter and how they should be managed. But lessons from chemical and radiation regulation also suggest we would be seriously mistaken in thinking that better measurement will resolve policy disputes. Ralph Nader, writing about the regulatory project at the time, was skeptical of the entire risk-assessment paradigm, calling it a movement toward “massive overcomplication and overabstraction” that tried to make precise something that cannot, by nature, be precise, and whose focus on ever-more-refined quantification could become a permanent reason to delay action on harms already obvious to the people experiencing them. Risk assessment, in his reading, simply relocated the political fight into a framework that looked technical but ultimately wasn’t.
AI risk management conversations have focused on the definition of risk being risk = hazard x probability. Risk communication expert Peter Sandman offers a different framework: risk = hazard + outrage, highlighting that the exercise of risk management must consider both, and that ignoring the signal he terms outrage corrodes the legitimacy of risk management as a whole:3
What the risk assessors mean by risk is not what anybody else means by risk. To a risk assessor, risk is a multiplication of magnitude — how bad is it when it happens? — and probability — how likely is it to happen? Let’s call that “hazard,” and let’s call what the public means by risk “outrage.” In outrage terms, a big outrage is a big risk. The experts focus on hazard and ignore outrage, whereas the public focuses on outrage and ignores hazard.
He went on:
How can we achieve technically optimal policies? By merging technical with political and psychological concerns at the beginning. Quantitative risk assessment is a good thing; it is not perfect, but it is better than nothing. It is the best tool we have for deciding how big a hazard is. But if we use it as an excuse for paying less attention to the public, less attention to outrage, then the outrage is going to increase. The gap between the public’s and experts’ concerns will increase, and quantitative risk assessment will be discredited in the process. An autocratic, unresponsive, untrustworthy risk manager is still a tyrant, even if he or she has better data.
A great deal of current AI concerns — discriminatory hiring tools, invasive facial recognition by police, AI companions targeted at children, generative tools fabricating intimate images of named individuals — sits in the high-outrage quadrant. Many of these have substantial evidence that has been pushed aside given the narrowness of the underlying technical systems, while some have weaker quantified evidence than the technical safety community typically looks for. Recent efforts to include kids safety in AI policy efforts suggest that some are coming to realize this, but that is just the tip of the iceberg. Deprioritizing concerns that people see as salient in their own lives is exactly the move that hollowed out trust in environmental health regulation across multiple decades. At the same time, outrage that is untethered to reality can suck up oxygen and resources truly needed to address harms for which scientific evidence is clear even if attention to those harms is low, which would also be a problematic failure model.
Measurement is policy, and other lessons
For AI policy, the history of environmental health regulation offers a few important lessons.
First, we’re in the early years of what will almost certainly be a long-running fight about what AI risks matter, and how those risks get measured. Private actors, poised to reap massive windfalls, will be incentivized to capture these conversations while suppressing the evidence they generate. Unless public interest actors invest seriously in the methodological terrain, the AI industry’s deep familiarity with the systems of interest and their technical details will, by default, shape choices about measurement. And like scientific research about chemicals, evidence collected inside of the companies who stand to profit from their own innovations is at risk of not becoming public for years. As evidence does emerge, the fight will become a fight over what counts as measurement, who is qualified to do it, and which methodological defaults sit underneath formal processes (J. Nathan Matias has done great work mapping out how these dynamics will play out and working to empower affected communities to broaden the frame of what evidence counts). This fight may continue for decades; we should plan for it to continue indefinitely.
Second, we need to expand attention beyond acute, immediately tractable harms from AI to cumulative ones that may not be immediately visible, but that nevertheless will still pose substantial, concrete harms. As with chemical and radiological exposure, tangible harms can build up steadily, with more exposure leading to higher likelihood of harm, or a more acute form of it. An automated hiring decision that explicitly rejects a candidate is easy to observe; a candidate ranking system that consistently ranks certain candidates lower accumulates over time in a way that has just as serious an effect, but in a far less visible way. And like chemical exposure and subsequent health problems, the causal relationship between exposure to AI and resulting harm is likely to be complex, so lessons from models like toxic torts (and how it has navigated questions about statutory presumptions of causation, burden shifting, and the role of scientific evidence) and Superfund programs (to remediate large-scale exposures) may be instructive.
Third, keeping risk assessment and risk management distinct is critical for both earning trust and ensuring accountability. Accepting risk means putting real people in harm’s way, and those choices must be made deliberately and with sound oversight. Risk management should make AI risk-acceptability decisions explicit, which means separating those decisions from methodological ones. Trust in AI is already low, and outrage is high. Separating risk assessment and risk management for AI would force policymakers and businesses to make their risk appetite publicly legible, enabling more robust public participation in conversations about whether the people and organizations who are making risk decisions are putting the public at more risk than it is willing to bear.
Finally, we can’t let the legitimate gaps in measurement methods distract from signals we already have about how AI is impacting people’s lives. Frontier AI systems pose some novel risk, which may indeed require new approaches to measurement, but narrow AI systems haven’t gone away, and they are already harming people with all too few guardrails.
Measuring the performance and outcomes of narrower systems is a tractable problem, but motivated actors have leveraged the complexity of frontier AI systems to divert attention to unsolved problems rather than the ones that could already be tackled, given sufficient resources. The playbook of sowing doubt about scientific consensus was honed in the context of public health toxicology first by the tobacco industry and then around climate change, and has already been applied to policy debates about tech-driven harms. Measurement will always involve some amount of uncertainty, and we need to figure out how to hold that reality alongside the possibility that some actors will try to exploit that uncertainty to advocate for inaction.
The environmental health analogy is not perfect. The set of chemicals to regulate is reasonably finite, and the physics of radiation exposure is stable, while AI models continue to proliferate, and minor tweaks can meaningfully change the safety profiles of those models. The nature of many AI systems, particularly generative ones, as information synthesis and content production machines means that constraining their deployment carries free-expression implications that don’t arise for industrial chemicals. Such real differences will complicate any efforts to port this regulatory tradition wholesale. But the contours of what environmental health regulation has learned about measurement under uncertainty, about competing risk models, about the politics of methodology do translate. AI policy will be living with these lessons either way. Better to learn them by example than the hard way.
Thanks to Kenneth Bogen, Kevin Bankston, and Ami Fields-Meyer for their helpful feedback and expertise.
If you don’t have time to read up on the entire scientific and policy history, the legal thriller Dark Waters gives an excellent and depressing window into the story.
Risk Assessment in the Federal Government: Managing the Process, National Research Council, 1983. Washington, DC: The National Academies Press. https://doi.org/10.17226/366.
Peter Sandman, “Definitions of Risk: Managing the Outrage, Not Just the Hazard,” Regulating Risk: The Science and Politics of Risk, Thomas A. Burke et al., editors, National Safety Council / ILSI Risk Science Institute, 1993.




Fantastic piece, Miranda. Thanks for this.