Timeline of experiment design
This is a timeline of experiment design, attempting to describe significant and illustrative events in the history of the field.
Sample questions
The following are some interesting questions that can be answered by reading this timeline:
- What landmark experiments illustrate the evolution of experimental design and its methodological pitfalls?
- Sort the full timeline by "Event type" and look for the group of rows with value "Experiment". You will see landmark experiments spanning medicine, physics, psychology, ecology, and economics, along with the methodological lessons each one illustrates.
- What are the core statistical and methodological concepts that underpin modern experimental design, and when were they introduced?
- Sort the full timeline by "Event type" and look for the group of rows with value "Concept". You will see the introduction of foundational statistical ideas as well as more recent methodological concerns.
- Which individuals and institutions pioneered specific experimental design methods, and what techniques did they introduce?
- Sort the full timeline by "Event type" and look for the group of rows with value "Method". You will see the introduction of a range of design and analysis techniques, from classical randomization schemes to modern computational approaches.
- How have ethical and regulatory frameworks for human experimentation developed over time?
- Sort the full timeline by "Event type" and look for the group of rows with value "Policy". You will see the emergence of research-ethics codes and regulatory requirements shaping how experiments involving human subjects are conducted and reported.
- Which organizations and institutions have shaped the field of experimental design?
- Sort the full timeline by "Event type" and look for the group of rows with value "Organization". You will see the founding of academic departments, professional bodies, and international collaborative networks central to the field.
- Other event types not covered by a dedicated question above include Statistical method, Publication, Concept development, Ethical oversight, Recognition, and Literature.
Big picture
| Time period | Development summary | More details |
|---|---|---|
| Ancient Times – 1700s | Early foundations | Experimental thinking has ancient roots, notably in medicine and agriculture. One of the earliest recorded episodes resembling a clinical trial dates to around 500 BCE, when King Nebuchadnezzar compares a meat-and-wine diet against one of beans and water, observing the latter group appears healthier.[1] Greek physicians like Hippocrates[2] and Galen emphasize observation and rudimentary comparative methods. However, formal experimentation remains rare. The Scientific Revolution in the 1600s shifts focus from authority to empirical investigation. Francis Bacon promotes inductive reasoning and systematic observation, laying philosophical groundwork for modern science.[3] Nonetheless, experiments lack standardized design or statistical rigor. Scientists like Galileo[4] and Boyle[5] use controlled observations, but their methods are more qualitative than quantitative. The period ends with the emergence of probability theory by Pascal[6], Fermat[7], and Bernoulli[8], which begins to offer tools later essential for experimental statistics. |
| 1700s – Late 1800s | Classical period | The Age of Enlightenment brings advances in both scientific methodology and mathematical statistics.[9] Controlled experiments become more common in chemistry, physics, and agriculture. Notably, James Lind conducts one of the first controlled clinical trials in 1747, testing treatments for scurvy.[10] In parallel, Carl Linnaeus systematizes biological classification, influencing observational rigor.[11] Pierre-Simon Laplace and Carl Friedrich Gauss develop early statistical tools such as the normal distribution and least squares estimation, crucial for data analysis. Agricultural trials, especially in Europe, begin to use basic comparison and replication. However, randomness, replication, and formal control are not yet standardized. This period sets the stage for modern experimentation by linking empirical testing with early statistical theory, but experiments still lack formal design structures and reproducibility standards. The idea of a placebo effect—a therapeutic outcome derived from an inert treatment— is already discussed in 18th century psychology[12] |
| Early 1900s – 1950s | Formalization and statistics | This period marks the birth of modern experimental design. Ronald A. Fisher revolutionizes the field in the 1920s and 1930s with work at Rothamsted Experimental Station in England.[13] Fisher introduces core principles: randomization, replication, and blocking. He develops the analysis of variance (ANOVA) and emphasizes the role of statistical inference in drawing conclusions. His 1935 book, The Design of Experiments, would remain foundational. Experimental design becomes central to agriculture[14][14][15], biology, and later psychology and social sciences. Jerzy Neyman and Egon Pearson introduce hypothesis testing frameworks, further refining experimental methodology.[16] This era institutionalizes statistics as an essential scientific tool and formalizes the structure of experiments, enabling reproducibility and greater objectivity in scientific inference. The need to blind researchers becomes widely recognized in the mid-century.[17] |
| 1960s – present | Modern and computational era | Post-1950s, experimental design expand into diverse fields: psychology, economics, medicine, and engineering. With the rise of computers, simulation, factorial designs, and response surface methodology become widespread. Clinical trials adopt strict protocols, often randomized and double-blind, guided by ethical oversight. In social sciences, randomized controlled trials (RCTs) gain ground, supported by behavioral economics and development studies. The Bayesian approach gains traction, offering alternative inference frameworks. From the 1990s, machine learning and A/B testing transform experimental design in tech industries, allowing real-time, large-scale experimentation. Today, experimental design is deeply intertwined with data science, AI, and evidence-based policy, continually evolving to address high-dimensional data, ethical concerns, and reproducibility challenges across disciplines. |
| Time period | Development summary | More details | |
|---|---|---|---|
| 1910s | Concept | Ronald A. Fisher develops the concept of variance as a measure of population variability, laying the foundation for later statistical methods such as analysis of variance (ANOVA).[18] | United Kingdom |
| 1920s | Method | Ronald Fisher develops randomized block design and formal principles of blocking to control variation in experiments. | United Kingdom |
| 1920s | Method | Fisher introduces key principles of modern experimental design, including confounding, randomization, replication, blocking, Latin square designs, and other factorial designs.[18] | United Kingdom |
| 1920s | Method | analysis of variance and maximum likelihood estimation are developed as major statistical tools for analyzing experimental data.[18] | United Kingdom |
| 1920s | Organization | The first academic statistics department is established at Iowa State University under George Snedecor, helping institutionalize statistical experimental methods.[18] | United States |
| 1920s | Method | George Snedecor develops the F-test and F-distribution, which become central tools for decisions in ANOVA-based experimental designs.[18] | United States |
| 1930s | Concept | The size of an experiment is recognized as a key design decision, with sample size linked to precision, significance, and statistical power.[19] | Global |
| 1930s | Concept | The advantages of factorial experimentation over one-factor-at-a-time methods are articulated, especially its economy of resources and ability to estimate interrelationships among factors.[19] | Global |
| 1930s | Field development | Early applications of statistical experimental design begin to appear in industrial settings, extending beyond agriculture.[20] | Global |
| 1930s | Method | Randomization is formalized as the use of random numbers or equivalent devices to assign treatments and reduce bias from extraneous variation.[19] | Global |
| 1930s | Concept | Experimental error is formally distinguished as a major obstacle to precision, arising from extraneous sources of variation beyond the treatments themselves.[19] | Global |
| 1930s | Concept | Bias is distinguished from random error, emphasizing that systematic error can mislead conclusions and often cannot be detected through statistical analysis alone.[19] | Global |
| 1930s | Method | double-blind experiments and placebo controls are recognized as important safeguards against bias in experiments involving subjective or clinical judgments.[19] | Global |
| 1930s | Method | Refinements of experimental technique, such as practice runs, clearer instructions, environmental control, and improved measurement instruments, are emphasized as ways to reduce experimental error.[19] | Global |
| 1930s | Method | blocking is presented as a method for increasing experimental precision by balancing treatments across subjects or units with similar initial characteristics.[19] | Global |
| 1930s | Method | matched pairs design and matched groups designs are used with human subjects as forms of blocking analogous to blocks in agricultural experiments.[19] | Global |
| 1930s | Concept | Core principles of experimental design—randomization, replication, and blocking—are established as fundamental techniques to reduce bias and improve precision in experiments.[21] | United Kingdom |
| 1930s | Method | analysis of covariance is introduced as a more accurate statistical adjustment for initial differences when blocking is not feasible or when further precision is needed.[19] | Global |
| 1930s | Method | factorial experimentation is formalized as a way to investigate several factors simultaneously, improving efficiency and enabling study of interactions among variables.[19] | Global |
| 1940s | Method | George Box and collaborators develop early central composite design methods for response surface optimization. | United Kingdom |
| 1940s | Method | sequential analysis emerges during World War II to improve efficiency in military experimentation and decision-making.[14] | United States |
| 1940s | Method | Development of orthogonal designs and Latin squares as structured experimental layouts for controlling variation in experiments.[14] | United Kingdom |
| 1950s | Method | W. Edwards Deming promotes statistical quality control in Japan, initiating the quality revolution and widespread industrial use of experimental design.[14] | Japan |
| 1950s | Method | Statistical methods and experimental design spread rapidly through the medical literature, extending techniques developed for agriculture into biomedical research.[18] | Global |
| 1950s | Method | Factorial experimental designs become widely used in industrial experimentation to study multiple variables simultaneously | United States |
| 1950s | Method | Control groups are emphasized as essential in evaluatory studies for distinguishing treatment effects from other changes over time.[19] | Global |
| 1950s–1970s | Field development | Statistical experimental design expands from agriculture into industrial applications, particularly in chemical and process industries.[22] | Global |
| 1950s–1980s | Field development | Response surface methods and statistical design techniques spread across chemical and process industries, especially in research and development.[20] | Global |
| Mid-20th century | Method | Standardized procedural steps for experimental design are established, including problem definition, variable selection, factor identification, experimental execution, and statistical analysis.[21] | Global} |
| Mid-20th century | Method | analysis of variance (ANOVA) becomes the central statistical framework for analyzing experimental data and guiding experimental design decisions.[21] | Global |
| 1960s | Method | Japanese industries adopt experimental design and statistical quality control, leading to major improvements in manufacturing quality and global competitiveness.[14] | Japan |
| 1960s | Experiment | randomized controlled trial becomes the gold standard in clinical research, replacing anecdotal medical evidence with controlled experimental methods.[14] | United States |
| 1960s–1970s | Concept | Industrial experiments are recognized as distinct due to immediacy and sequential learning, enabling rapid iteration and adaptive experimentation.[20] | Global |
| 1970s | Method | Multi center clinical trials become standard, enabling large-scale randomized evaluation across diverse populations | Global [23] |
| 1970s–1980s | Method | Genichi Taguchi develops robust parameter design methods to reduce variability and improve product quality.[14] | Japan |
| 1970s–1990 | Method | Quality improvement methodologies such as Total Quality Management (TQM) and Continuous Quality Improvement (CQI) integrate experimental design into industrial processes.[14] | Global |
| Late 1970s | Method | Genichi Taguchi promotes robust parameter design, emphasizing quality improvement and reduction of variability under noise conditions using orthogonal arrays and fractional factorial designs.[24][20] | Japan |
| Late 1970s–1980s | Field development | Growing interest in quality improvement in Western industry leads to broader adoption of experimental design methods.[20] | Global |
| 1980s | Method | Quality engineering approaches emphasize robustness and variance reduction in manufacturing processes | Japan [25] |
| 1980s | Method | Meta analysis emerges as a method to systematically combine results from multiple experiments | Global |
| 1980s | Field development | Designed experiments become widely adopted in manufacturing industries (automotive, aerospace, electronics), driven by quality improvement initiatives.[22] | Global |
| 1980s | Method | Large-scale randomized trials such as the ISIS heart studies demonstrate the feasibility of very large clinical experiments.[26] | Global |
| Late 1980s | Concept | Critical evaluation of Taguchi methods leads to refinement of statistical foundations and development of alternative robust design methodologies.[22] | Global |
| 1990s | Method | Six Sigma integrates statistical experimental design into industrial quality control and process improvement,[14] | United States |
| 1990s | Method | optimal design approaches gain prominence, focusing on maximizing information and efficiency under resource constraints, supported by advances in computing.[22] | Global |
| 1990s | Field development | Designed experiments expand into new domains including services, finance, and government operations.[22] | Global |
| 1990s | Method | The modern era of experimental design begins, driven by globalization and economic competition, expanding DOE across industries and services.[14] | Global |
| 1990s–present | Field development | The Optimal Design Era emerges, characterized by widespread use of computational algorithms, software tools, and expansion of DOE across diverse sectors.[22] | Global |
| 2000s | Method | Online randomized controlled experiments (A/B testing) become standard in internet companies for product and interface optimization, enabling continuous data-driven decision making.[27] | United States |
| 2000s | Method | Bradley Jones introduces Bayesian D-optimal design approaches, incorporating prior knowledge into experimental planning.[22] | United States |
| 2000s | Method | JMP (software)'s Custom Design platform, part of SAS Institute's design-of-experiments software, enables flexible, computer-generated optimal experimental designs tailored to specific problems; constraint-handling features for the platform are progressively extended starting with JMP 6.[28] | United States |
| 2000s | Method | Advances in supersaturated designs and split-plot designs improve efficiency in experiments with many factors and constraints.[22] | United States |
| 2000s–present | Method | Development of group-orthogonal supersaturated designs (GO-SSDs) improves factor screening when variables exceed experimental runs.[22] | United States |
| 2010s | Method | Continuous A/B testing systems enable rapid, large-scale experimentation in digital platforms, allowing ongoing optimization based on user behavior data.[29] | Global |
| Late 20th century | Application | Experimental design expands beyond agriculture into engineering, business, and scientific research, becoming a general-purpose methodology.[21] | Global |
Full timeline
| Year | Event type | Details | Geographical location |
|---|---|---|---|
| ca. 500 BCE | Experiment | According to the Book of Daniel, Daniel and his companions request a ten-day trial in which they consume only vegetables and water, while other youths at the royal court continue eating King Nebuchadnezzar's rich diet; the results of the two regimens are compared, and the steward acts based on the observable outcome. Several scholars of ancient literature and the history of science have cited this account as one of the earliest descriptions of a controlled comparative test, and biblical historian David C. Lindberg specifically includes it among ancient examples where a trial is deliberately arranged to evaluate competing conditions.[30][31][32] | Babylon (Mesopotamia) |
| ca. 500 BCE | Experiment | According to the Book of Daniel, Daniel and his companions request a ten-day trial in which they consume only vegetables and water, while other youths at the royal court continue eating the king's rich diet; the results of the two regimens are compared, and the steward acts based on the observable outcome. Several scholars of ancient literature and the history of science have cited this account as one of the earliest descriptions of a controlled comparative test, and biblical historian David C. Lindberg specifically includes it among ancient examples where a trial is deliberately arranged to evaluate competing conditions.[33][34][35] | Babylon (Mesopotamia) |
| ca. 400 BCE | Concept | Hippocrates promotes systematic observation and naturalistic explanations in medicine, emphasizing empirical study over supernatural interpretations.[1] | Greece |
| ca. 587 AD | Concept | Indian astronomer-mathematician Varahamihira, in his text Brhat Samhita, describes what is considered one of the earliest datable applications of combinatorial design: a method for making perfumes by selecting 4 substances from 16 available substances using a magic square arrangement.[36] | India |
| c. 1021 | Method | Arab polymath Ibn al-Haytham completes the Book of Optics, applying a methodical, inductive-experimental approach to optical and mathematical problems inherited from Ptolemy — emphasizing self-criticality, reliance on visible experimental results over inherited authority, and rigorous skepticism toward earlier scholars' conclusions. He explicitly instructs that a researcher should "make himself an enemy of all that he reads" and interrogate it from every angle, while remaining alert to his own susceptibility to prejudice and leniency. Historians of science regard Ibn al-Haytham as one of the first scholars to systematically apply controlled, inductive experimentation to derive knowledge, prefiguring elements of the modern scientific method by roughly six centuries.[37] | Egypt |
| 1025 AD | Concept | Persian polymath Avicenna completes The Canon of Medicine, a major medical encyclopedia synthesizing Greco-Roman and Islamic knowledge. The work describes diseases, treatments, and hundreds of drugs, and outlines principles for experimentally testing medicines, including reproducibility of results, shaping medical education in both the Islamic world and Europe for centuries.[38] | Persia |
| 1620 | Concept | Francis Bacon publishes Novum Organum, proposing a new empirical approach to scientific inquiry. He argues that knowledge should be based on systematic observation and inductive reasoning rather than tradition or pure logic, helping establish methodological foundations for the modern scientific method and experimental science.[39] | England |
| 1648 | Concept | Flemish physician and chemist Jan Baptist van Helmont, in his posthumously published Ortus Medicinae, proposes what historians later identify as the first documented design for a randomized comparative trial: dividing 200–500 poor patients suffering from fevers or pleurisy into two groups by lot, treating one group with the standard Galenic methods of bloodletting and purging and the other without them, and comparing the resulting death counts between groups. There is no evidence the trial was ever actually conducted, but the proposal anticipates by nearly three centuries the logic of random allocation between a treatment and a comparison group.[40] | Flanders |
| 1676 | Experiment | Danish astronomer Ole Rømer demonstrates that light does not travel instantaneously but has a measurable finite speed, by observing that the apparent timing of Jupiter's moons' eclipses was delayed when Jupiter was farther from Earth compared to when it was closer — an early landmark example of a natural experiment, relying on observation of a system as it naturally occurs rather than direct manipulation of variables.[41] | Denmark |
| 1683 | Method | Robert Boyle solidifies his reputation as a pioneer of modern experimental science by publishing New Experiments and Observations Touching Cold. During this period, Boyle and his contemporaries, including Robert Hooke, emphasize that chemical understanding must be grounded in controlled experimentation, precise documentation of apparatus, and the replication of results.[42] | England |
| 1700 | Concept | Korean mathematician Choi Seok-jeong is the first to publish an example of Latin squares of order nine, in order to construct a magic square, predating Leonhard Euler by 67 years.[43] Latin squares are used in combinatorics and in experimental design.[44] | Korea |
| 1710 | Concept | John Arbuthnot publishes "An Argument for Divine Providence, taken from the constant regularity observed in the births of both sexes" in Philosophical Transactions of the Royal Society, examining London birth records for each of 82 years from 1629 to 1710 and finding that male births exceeded female births in every single year. Applying what becomes known as the sign test, Arbuthnot calculates that the probability of this occurring by chance alone (if male and female births were equally likely) is 0.582 — an astronomically small figure — and concludes the pattern reflects design rather than chance. The paper is widely credited as the earliest recorded use of statistical hypothesis testing, predating the formal apparatus of significance testing by over two centuries.[45] | England |
| 1713 | Concept | Jacob Bernoulli formulates the law of large numbers, establishing convergence of sample averages to expected values. | Switzerland |
| 1747 | Experiment | Scottish doctor James Lind conducts one of the earliest controlled clinical trials when investigating the efficacy of citrus fruit in cases of scurvy, dividing twelve scurvy patients, whose "cases were as similar as I could have them", into six pairs — without randomization, but with an early attempt at controlling for similar starting conditions. Each pair is given a different remedy. According to Lind's 1753 Treatise on the Scurvy in Three Parts Containing an Inquiry into the Nature, Causes, and Cure of the Disease, Together with a Critical and Chronological View of what has been Published of the Subject, the remedies were: one quart of cider per day, twenty-five drops of elixir vitriol (sulfuric acid) three times a day, two spoonfuls of vinegar three times a day, a course of sea-water (half a pint every day), two oranges and one lemon each day, and electuary, (a mixture containing garlic, mustard, balsam of Peru, and myrrh).[46][21] Lind would note that the pair who had been given the oranges and lemons were so restored to health within six days of treatment that one of them returned to duty, and the other was well enough to attend the rest of the sick.[46] | United Kingdom |
| 1772 | Concept | Scottish physician William Cullen, then the most influential medical lecturer in the Anglo-American world at the University of Edinburgh, becomes the first person known to use the word "placebo" in a medical sense, in his manuscript clinical lectures. Cullen prescribes a "placebo" — typically a weak but genuinely active substance such as mustard or Dover's powder — to a patient he considers "absolutely incurable," stating he does so "in pure placebo" while trying, "even in employing placebos," "to give what would have a tendency to be of use to the patient." For Cullen, a placebo is defined less by its chemical inertness than by the physician's lack of curative intent — what later historians term an "active placebo," distinct from the inert (bread- or sugar-pill) sense the term would come to carry by the early 20th century. The word itself derives from the Latin Placebo Domino ("I shall please the Lord," Psalm 116:9), used since the medieval period to name the Vespers for the Dead, and had by Chaucer's time already acquired a secular, pejorative sense — a flatterer who "sings placebos."[47] | Scotland |
| 1778 | Concept | Pierre-Simon Laplace examines birth statistics from almost half a million births in "Mémoire sur les probabilités," finding a consistent excess of male births over female births and concluding through probabilistic calculation that the excess is a real, though unexplained, effect rather than a product of chance — extending the sex-ratio hypothesis-testing tradition begun by John Arbuthnot in 1710 to a substantially larger dataset and a more developed probabilistic apparatus.[48] | France |
| 1782 | Concept | Leonhard Euler publishes "Recherches sur une nouvelle espèce de quarrés magiques" in the Verhandelingen of the Zeeland Society of Sciences (first presented to the Imperial Academy of Sciences of St. Petersburg in 1779), using Latin characters as symbols in his arrays and thereby giving rise to the name "Latin square." Though Korean mathematician Choi Seok-jeong had published an example of order-nine Latin squares 67 years earlier, Euler's work begins the general mathematical theory of Latin squares and their combinatorial properties.[49] | Russia |
| 1780 | Concept | Franz Mesmer, then already well known in Paris, refuses a request from the Paris Faculty of Medicine's Joseph-Marie-François de Lassone to have his therapeutic method scrutinized through new cures on unfamiliar patients, arguing instead that his prior "cures" should be taken as an objective matter of record — a position that both 1784 Royal Commissions would later reject in designing their own investigations. Separately, Mesmer proposed principles for objectively testing his method that anticipate elements of controlled experimental design.[50] | France |
| 1784 | Experiment | The first blinded experiment is conducted by the French Academy of Sciences to investigate the claims of mesmerism as proposed by Franz Mesmer. In the experiment, researchers blindfolded mesmerists and asked them to identify objects that the experimenters had previously filled with "vital fluid". The subjects are unable to do so.[51] | France |
| 1784 | Experiment | Two independent French Royal Commissions — a nine-member "Franklin Commission" (four physicians from the Paris Faculty of Medicine and five scientists from the Royal Academy of Sciences, including Benjamin Franklin and Antoine Lavoisier) and a five-member "Society Commission" from the Royal Society of Medicine — are appointed by Louis XVI to investigate Charles d'Eslon's claims for the physical existence of "animal magnetism." Lavoisier designs the Franklin Commission's investigative protocol: rather than assessing long-term cures, the commissioners test the "momentary" physiological effects of magnetization under conditions systematically varying whether subjects are genuinely or falsely told they are being magnetized, and whether investigators and subjects are blindfolded. In one representative test, a "sensitive" subject who believed he had been led to a magnetized tree fainted at the foot of the wrong tree; in another, a subject who drank ordinary water believing it magnetized displayed magnetic symptoms. Both Commissions conclude that d'Eslon's magnetic fluid does not exist and that all observed effects are attributable to touch, imagination, and imitation. The Franklin Commission's report, presented 11 August 1784 and printed in at least 20,000 copies, is widely regarded by historians of science as one of the earliest classic examples of a systematic controlled trial, notable in particular for its use of literal blindfolding of both investigators and subjects.[52][53] | France |
| 1785 | Publication | The word "placebo" appears in medical print for the first time, in the second edition of George Motherby's New Medical Dictionary, which defines it as "a commonplace method or medicine"; the 1775 first edition had contained no such entry. Historian of the placebo Arthur Shapiro would later call the reasons for the term's entry into medicine "largely unknown" — a gap subsequent research traces to William Cullen's clinical teaching over the preceding decade.[47] | England |
| 1785 | Publication | The word "placebo" appears in medical print for the first time, in the second edition of George Motherby's New Medical Dictionary, which defines it as "a commonplace method or medicine"; the 1775 first edition had contained no such entry. The word had entered English via the Vulgate's rendering of Psalm 116:9 as "Placebo Domino" ("I shall please the Lord"), likely chosen by Jerome for its metre rather than as an accurate translation of the Hebrew; by the 13th century "Placebo" named the Vespers for the Dead, and by Chaucer's time had acquired the pejorative sense of a flattering sycophant. Historian of the placebo Arthur Shapiro would later call the reasons for the term's entry into medicine "largely unknown" — a gap subsequent research traces to William Cullen's clinical teaching over the preceding decade.[47][54] | England |
| 1796 | Experiment | English physician Edward Jenner inoculates eight-year-old James Phipps with cowpox matter taken from dairymaid Sarah Nelms's lesions on 14 May 1796, then deliberately exposes him to smallpox matter that July; Phipps develops no disease, and Jenner concludes the boy is protected. Jenner expands his findings into An Inquiry into the Causes and Effects of the Variolae Vaccinae, privately published in 1798 after the Royal Society rejected his initial 1797 submission. A landmark early medical experiment, it tests a single subject without any untreated comparison group to rule out chance or an unrelated cause for the observed protection — the same limitation found in Lady Mary Wortley Montagu's earlier 1721 variolation trial on six condemned prisoners at Newgate, whose survival and subsequent immunity likewise had no control group to benchmark against.[55] | England |
| 1798 | Statistical method | German mathematician Carl Friedrich Gauss develops the mathematical foundations of the method of least squares. Years later, the method would enable astronomers to predict the orbit of the asteroid Ceres after its discovery by Giuseppe Piazzi and help Franz Xaver von Zach successfully relocate it.[56] | Germany |
| 1799 | Concept | English physician John Haygarth, working in Bath, tests the popular but ultimately fraudulent "Perkins tractors" — a pair of metal rods invented by Elisha Perkins and marketed as curing pain and disease through "animal magnetism" — by comparing their effects against plain wooden rods disguised to look identical. Finding the wooden sham tractors work just as well as the metal originals, Haygarth demonstrates the "cure" comes from patient expectation rather than the device itself, in what historians regard as one of the earliest uses of a placebo control in a single-blind clinical trial.[57] | England |
| 1811 | Concept | The word "placebo" is defined in Hooper's Medical Dictionary as "any medicine adapted more to please than benefit the patient" — a definition markedly closer to the modern sense of the term than Motherby's 1785 entry ("a commonplace method or medicine"), and one that explicitly separates a placebo's pleasing effect from any genuine therapeutic benefit for the first time in print. Physician and pharmacologist Jeff Aronson later speculates the shift may reflect a turn-of-the-century reinterpretation of the word's derivation, a form of medical snobbery toward "popular" remedies, or a reversion to the term's older, sycophantic connotation.[54] | England |
| 1815 | Concept | An article on optimal designs for polynomial regression is published by Joseph Diaz Gergonne.[58] | France |
| 1817 | Experiment | The first blinded experiment recorded outside of a scientific setting compares the musical quality of a Stradivarius violin to one with a guitar-like design. A violinist plays each instrument while a committee of scientists and musicians listen from another room so as to avoid prejudice.[59][60] | France |
| 1827 | Method | Pierre-Simon Laplace uses least squares methods to address analysis of variance problems regarding measurements of atmospheric tides.[61] | France |
| 1835 | Experiment | An early example of a double-blind protocol is the Nuremberg salt test performed by Friedrich Wilhelm von Hoven, Nuremberg's highest-ranking public health official.[62] | Germany |
| 1835 | Concept | Belgian statistician Adolphe Quetelet introduces the concept of the “average man” (l’homme moyen), arguing that human physical and social traits follow statistical distributions. By applying probability and quantitative analysis to crime, mortality, and demographics, he helps establish statistical approaches in sociology, demography, and the emerging social sciences.[63] | Belgium |
| 1843 | Concept | Irish doctor James Henry proposes principles for conducting a controlled trial comparing cold-water therapy against sulphur treatment for scabies — an early articulation of controlled-trial methodology predating the formalization of randomized experimental design by nearly a century.[64] | Ireland |
| 1844 | Concept | Wesley S. B. Woolhouse poses the problem of what becomes known as Steiner triple systems as Prize Question #1733 in the Lady's and Gentlemen's Diary — asking for a systematic partition of elements into triples such that every pair of elements appears together in exactly one triple.[65] | England |
| 1847 | Method | English clergyman and mathematician Thomas Kirkman solves Wesley S. B. Woolhouse's 1844 prize problem in "On a Problem in Combinations," published in The Cambridge and Dublin Mathematical Journal — the first solution to what becomes known as the Steiner triple system problem, three years before Kirkman posed his own, more elaborate variant asking for resolvable systems (later known as Kirkman's schoolgirl problem).[66] | England |
| 1850 | Concept | Thomas Kirkman poses the problem in the Lady's and Gentleman's Diary of 1850: "Fifteen young ladies of a school walk out three abreast for seven days in succession: it is required to arrange them daily so that no two shall walk abreast more than once." This becomes known as the Fifteen Schoolgirls Problem, or Kirkman's schoolgirl problem. A solution to this recreational puzzle is equivalent to a resolvable balanced incomplete block design with 15 points, block size 3, and λ = 1 — an early landmark in the theory of resolvable designs. Arthur Cayley published a solution first, followed by Kirkman's own (which he had, of course, already worked out before posing the question); James Joseph Sylvester also studied the problem and later disputed priority with Kirkman.[67] | England |
| 1853 | Concept | Jakob Steiner independently reintroduces triple systems in "Combinatorische Aufgabe," published in Crelle's Journal, unaware of Thomas Kirkman's 1847 solution to the same problem. Because Steiner's paper becomes more widely known within the mathematical community than Kirkman's earlier work, the systems come to be named Steiner systems in his honor rather than Kirkman's, despite Kirkman's priority.[68] | Germany |
| 1854 | Experiment | English physician John Snow investigates a cholera outbreak that kills over 500 people in ten days in London's Soho district, mapping cholera deaths against residence and finding that victims were disproportionately concentrated among those drinking from a shared pump on Broad Street. He also finds that brewery workers and workhouse residents nearby, who drew water from separate wells rather than the Broad Street pump, were largely spared. Snow persuades local officials to remove the pump's handle; the outbreak, already declining as residents fled the area, ceases entirely within days. The bacterial mechanism of cholera transmission would not be confirmed until Robert Koch's later work.[69] | England |
| 1860 | Method | German psychologist and physicist Gustav Theodor Fechner makes a groundbreaking contribution to the development of experimental psychology with the publication of Elements of Psychophysics. In this work, he seeks to test and justify the relationship between physical stimuli and the sensations they produce. Fechner proposes that mental experiences can be quantified by linking them to measurable physical changes, laying the foundation for psychophysics. Using experimental data, he formulates mathematical laws—most notably the Weber-Fechner law—that describe how perceived intensity varies with stimulus magnitude. His efforts help establish psychology as a quantitative science rooted in empirical observation.[70] | Germany |
| 1861 | Experiment | French chemist and microbiologist Louis Pasteur conducts controlled experiments demonstrating that microorganisms originate from airborne particles rather than spontaneous generation. In his memoir examining this doctrine, he shows that sterilized broth remain free of life unless exposed to contaminated air, helping establish the foundations of modern microbiology and germ theory.[71][72] | France |
| 1865 | Concept | French physiologist Claude Bernard publishes Introduction to the Study of Experimental Medicine, advocating that researchers should not know the hypothesis being tested while making observations — a recommendation that directly contradicted the prevailing Enlightenment-era view that valid scientific observation required a well-educated, informed observer.[73] | France |
| 1876 | Concept | American scientist Charles S. Peirce contributes the first English-language publication on an optimal design for regression models.[74] | United States |
| 1877 | Concept | Charles Sanders Peirce formalizes inquiry as a structured experimental process: hypotheses (abduction) generate testable predictions (deduction), which are evaluated through observation and experiment (induction). This integrates reasoning with empirical testing, establishing a cyclical, self-correcting model of experiment design focused on hypothesis testing and iterative refinement, developed across his "Illustrations of the Logic of Science" series.[75][76] | United States |
| 1879 | Organization | Experimental psychology emerges as a modern scientific discipline in Germany with the establishment of the first experimental laboratory by Wilhelm Wundt at the University of Leipzig. This marks a pivotal moment in the history of psychology, as Wundt seeks to separate psychology from philosophy by applying rigorous scientific methods. He introduces a mathematical and experimental approach to studying the human mind, emphasizing observation, measurement, and controlled experimentation. Wundt's work lays the foundation for psychology as an empirical science, influencing future researchers and schools of thought.[70] | Germany |
| 1879–1880 | Experiment | The Milwaukee Academy of Medicine conducts a trial testing a homeopathic remedy against a sugar-pill control, in which both the patients and the physician-experimenters (apart from those managing the code) are masked as to which substance had been administered — a design later historians describe in modern terms as "double-blind." The trial, reported in the Homoeopathic Times as the "Final report of the Milwaukee test of the thirtieth dilution," emerges from the broader 19th-century dispute between homeopathy and orthodox medicine, in which both sides increasingly turned to masked assessment to adjudicate rival therapeutic claims.[77][78] | United States |
| 1882 | Concept | In his published lecture at Johns Hopkins University, Peirce introduces experimental design with these words:
|
United States |
| 1884 | Organization | Frederick Akbar Mahomed, who worked at Guy's Hospital in London and separated chronic nephritis with secondary hypertension from what would later be termed essential hypertension, founds the Collective Investigation Record for the British Medical Association — an organization collecting data from physicians practicing outside the hospital setting, considered the precursor of modern collaborative, multi-site clinical trials.[80] | England |
| 1885 | Concept | Analysis of variance. An eloquent non-mathematical explanation of the additive effects model becomes available.[81] | United Kingdom |
| 1887 | Experiment | Albert A. Michelson and Edward W. Morley conduct an interferometry experiment designed to detect Earth's motion through the hypothesized luminiferous aether by measuring expected variations in the speed of light in different directions. The experiment fails to detect the predicted aether drift, instead measuring a drift far too small to account for the theoretically expected effect and generally attributed to noise — becoming one of the most famous null results in the history of science. The unexpected absence of a positive result would go on to contribute significantly to the development of special relativity.[82] | United States |
| 1880s | Experiment | Charles Sanders Peirce and Joseph Jastrow introduce randomized experiments in the field of psychology.[83] | United States |
| 1894 | Concept | French psychologist Alfred Binet publishes "La psychologie de la prestidigitation" in the Revue des Deux Mondes, examining how magicians exploit blind spots in human attention and perception to perform their tricks — among the earliest documented scientific studies of magic (illusion) as a subject in its own right. An English translation, "Psychology of Prestidigitation," appears in 1896 in the Smithsonian Institution's annual report. The field of magic as a subject of scientific study would not be substantially revived until the early 21st century.[84][85] | France |
| 1897 | Experiment | Norman Triplett conducts one of the first social psychology experiments on cyclist performance.[86] | United States |
| 1900 | Method | The P-value is first formally introduced by Karl Pearson, in his Pearson's chi-squared test, using the chi-squared distribution and notated as capital P.[87] Since then, P-values would become the preferred method to summarize the results of medical articles.[88][89] | United Kingdom |
| 1900 | Concept | French mathematician Gaston Tarry proves that Leonhard Euler's 36 officers problem has no solution, confirming Euler's own 1782 conjecture that no pair of orthogonal Latin squares of order six exists. Tarry's proof, published in Comptes Rendus de l'Association Française pour l'Avancement des Sciences, works by exhaustively checking all essentially distinct Latin squares of order six — a set Tarry reduces to 9,408 cases — and confirming that no two of them are orthogonal to each other. The conjecture would later prove false for every other order of the form 4k+2 above 6, as R. C. Bose, S. S. Shrikhande, and E. T. Parker would show in 1959–1960.[90] | France |
| 1900 | Concept | The United States Army's Yellow Fever Commission, led by Walter Reed, conducts mosquito-transmission experiments on yellow fever at a camp in Cuba using American and Spanish volunteers. The study marks the first documented instance of test subjects signing a consent form before participating in a medical experiment, with researchers stating they wanted volunteers to understand the hazards involved beforehand. None of the volunteers died, and the experiments confirm that mosquitoes transmit the disease.[91] | Cuba |
| 1903 | Concept | American physician Richard Clarke Cabot concludes that the placebo should be avoided because it is deceptive.[92] | United States |
| 1904 | Concept | English psychologist Charles Spearman becomes the first psychologist to discuss common factor analysis, in his paper "General Intelligence, Objectively Determined and Measured." Spearman observes that schoolchildren's scores across a wide variety of seemingly unrelated subjects are positively correlated, leading him to postulate a single general mental ability, "g," underlying human cognitive performance — the origin of factor analysis as a statistical technique, though Spearman's paper provides few methodological details and is concerned only with single-factor models.[93][94] | England |
| 1907 | Experiment | German philosopher and psychologist Carl Stumpf, together with his assistant Oskar Pfungst, investigates the claims surrounding "Clever Hans," an Orlov Trotter horse whose owner, Wilhelm von Osten, claimed could perform arithmetic and other cognitive tasks. Ruling out deliberate fraud, Pfungst determines the horse answers correctly 89% of the time when the questioner knows the answer, but only 6% of the time when the questioner does not — and shows that questioners unconsciously tensed their posture and expression as the horse's taps approached the correct answer, providing an unintentional cue the horse had learned to use as a stopping signal. The episode becomes the classic illustration of the observer-expectancy effect and a foundational argument for blinded experiment the questioner or experimenter, not just the subject, from knowledge that could unconsciously bias results.[95] | Germany |
| 1907 | Experiment | The first study recorded to have a blinded researcher is conducted by W. H. R. Rivers and H. N. Webber to investigate the effects of caffeine.[96] | United Kingdom |
| 1908 | Method | British statistician William Sealy Gosset, working at Guinness, introduces the Student’s t-distribution to analyze small sample data when population variance is unknown. Publishing under the pseudonym “Student” in the journal Biometrika, he provides a statistical method widely used for inference with limited data.[97] | United Kingdom |
| 1908–1909 | Concept | Irish economist and statistician Francis Ysidro Edgeworth publishes a series of papers "On the Probable Errors of Frequency-Constants" that anticipate several elements of what would become known as Fisher information, over a decade before Fisher's own 1922 formal treatment — work later recognized by statisticians and historians of statistics as an important, underacknowledged precursor.[98] | Ireland |
| 1911–1918 | Experiment | German physician Adolf Bingel conducts a large-scale blinded, controlled comparison of diphtheria treatments, alternately assigning 937 patients between diphtheria antitoxin serum and plain horse serum; all patients and participating physicians other than Bingel himself remain unaware of each patient's allocation. Published in 1918 in the Deutsches Archiv für klinische Medizin, the trial becomes an influential early example of blinded assessment applied at substantial scale in mainstream (rather than "irregular") medicine, feeding into a wider German tradition of blind clinical assessment.[77] | Germany |
| 1918 | Concept | English statistician Ronald Fisher introduces the term variance and proposes its formal analysis in his article The Correlation Between Relatives on the Supposition of Mendelian Inheritance.[99] | United Kingdom |
| 1918 | Method | Kirstine Smith publishes a major study in Biometrika analyzing the statistical precision of polynomial regression estimates. She derives principles for choosing observation points that minimize estimation variance, introducing foundational methods for optimal experimental design in polynomial models and influencing later developments in statistical design theory.[100] | Denmark |
| 1918–1940s | Method | Ronald A. Fisher and collaborators establish the foundations of modern experimental design in agricultural research, introducing factorial designs and analysis of variance (ANOVA).[14] | United Kingdom |
| 1919 | Organization | Ronald A. Fisher is hired as a statistician at the Rothamsted Experimental Station, where his work on poorly designed agricultural data helps trigger the modern statistical approach to experimental design.[18] | United Kingdom |
| 1919 | Method | R. A. Fisher at the Rothamsted Experimental Station in England starts developing modern concepts of experimental design in the planning of agricultural field experiments.[101] | England |
| 1920 | Concept | English physician T.C. Graves is the first to define and discuss the "placebo effect" in print, in a commentary published in The Lancet on a case of hystero-epilepsy with delayed puberty. Graves describes "the placebo effects of drugs" as occurring in cases where "a real psychotherapeutic effect appears to have been produced" despite the absence of any specific pharmacological action — using the phrase decades before Henry K. Beecher's more widely cited 1955 paper of the same name.[102] | England |
| 1921 | Method | Ronald Fisher publishes his first application of the analysis of variance.[103] | United Kingdom |
| 1922 | Concept | Ronald Fisher formally defines what becomes known as Fisher information in "On the Mathematical Foundations of Theoretical Statistics," measuring the amount of information an observable random variable carries about an unknown parameter of the distribution generating it. Fisher information becomes foundational to the theory of maximum-likelihood estimation and to optimal design of experiments, where maximizing the Fisher information of a design corresponds to minimizing the variance of parameter estimates obtainable from it.[104] | United Kingdom |
| 1923 | Field development | The first randomization model is published in Polish by Jerzy Neyman.[105] | Poland |
| 1923 | Method | Ronald A. Fisher, together with Winifred Mackenzie, publishes "Studies in Crop Variation II," extending his 1921 "Studies in Crop Variation I" (which had partitioned the variation of a time series into annual and slow-deterioration components) to study yield variation across plots sown with different potato varieties and subjected to different fertiliser treatments — early applied work at Rothamsted Experimental Station that would feed directly into the formal analysis-of-variance framework Fisher developed over the following years.[106][107] | United Kingdom |
| 1924–1925 | Experiment | Political scientist Harold Gosnell conducts an early field experiment on voter participation in Chicago, testing whether nonpartisan mailings encouraging registration and voting increased turnout — widely regarded as one of the earliest field experiments in political science, using randomization outside a laboratory setting to study a real-world civic behavior.[108] | United States |
| 1925 | Publication | British statistician Ronald Fisher publishes Statistical Methods for Research Workers, developing tests of statistical significance, proposing the 0.05 threshold for hypothesis testing, and popularizing analysis of variance as a method for partitioning variation in experimental data. The book would later be identified as a seminal, and controversial, influence — some retrospective critics tracing aspects of the modern replication crisis in science to conventions it popularized.[109][110][111][112][113][22][114] | United Kingdom |
| 1926 | Publication | John Russell (agricultural scientist) publishes "Field Experiments: How They Are Made and What They Are", summarizing contemporary practices and principles of experimental design in agricultural research.[115] | United Kingdom |
| 1926 | Experiment | Janet Lane-Claypon publishes a retrospective investigation into breast cancer risk factors in the United Kingdom, widely regarded as the first major case–control study — comparing women with the disease against a control group without it to isolate differences in reproductive history. Her data provide the first epidemiologic evidence that low fertility (fewer pregnancies, later age at first birth, and shorter duration of breastfeeding) increases breast cancer risk, findings a 2010 statistical reanalysis would confirm as consistent with modern epidemiologic evidence. The study is replicated in the United States in 1931 by J.M. Wainwright.[116] | United Kingdom |
| 1926 | Concept | Ronald Fisher argues in "The Arrangement of Field Experiments" that studying multiple factors simultaneously in "complex" (factorial) designs is more efficient than the traditional one-factor-at-a-time approach, writing that "Nature... will best respond to a logical and carefully thought out questionnaire; indeed, if we ask her a single question, she will often refuse to answer until some other topic has been discussed" — arguing that a well-designed factorial experiment can determine the effects of several factors, and their interactions, using no more trials than would be needed to determine just one factor's effect alone.[117] | United Kingdom |
| 1926 | Concept | Soviet mathematician Sergei Natanovich Bernstein introduces the "blocks method" in probability theory, splitting a sample into blocks separated by smaller subblocks so the blocks can be treated as approximately independent — a technique later used to prove limit theorems for sums of dependent random variables and applied in extreme value theory. This is a distinct lineage from Fisher's contemporaneous statistical blocking in experimental design, converging on the same term for a different underlying idea.[118] | Soviet Union |
| 1930 | Concept | Torald Sollmann, head of the American Medical Association's Council on Pharmacy and Chemistry, advocates that clinical drug testing adopt the rigor of other scientific disciplines, explicitly using the terms "comparative method," "double-blind procedure," and "blind test" in doing so — among the earliest documented uses of "double-blind" terminology in American medicine, predating its wider adoption by German clinical pharmacology under Paul Martini two years later.[119] | United States |
| 1931–1935 | Method | American psychologist Louis Leon Thurstone develops common factor analysis with multiple factors across two papers in the early 1930s, summarized in his 1935 book The Vectors of Mind: Multiple-Factor Analysis for the Isolation of Primary Traits. Thurstone introduces several concepts still central to the technique, including communality, uniqueness, and factor rotation, and advocates for achieving "simple structure" in factor solutions through appropriate choice of rotation.[93][120][121] | United States |
| 1932 | Method | German clinical pharmacologist Paul Martini publishes Methodenlehre der Therapeutischen Untersuchung ("Methodology of Therapeutic Investigation"), codifying blind assessment as a formal methodological principle for evaluating treatments — consolidating the German tradition of blinded experimentation that had developed since Adolf Bingel's 1918 diphtheria trial, shortly before the Nazi period disrupts German academic medicine.[77] | Germany |
| 1932 | Experiment | American physician Harry Gold, an early and influential advocate for masked assessment in American clinical pharmacology, tests xanthine against a lactose placebo in a single-blind trial for cardiac pain — a milestone experiment in the American adoption of placebo-controlled methodology, conducted the same year Paul Martini codifies double-blind testing in Germany.[119] | United States |
| 1933 | Concept | Jerzy Neyman and Egon Pearson publish "On the Problem of the Most Efficient Tests of Statistical Hypotheses" in Philosophical Transactions of the Royal Society A, part of a series of papers begun in 1928 that formalizes the statistical hypothesis test as a proposed refinement of Ronald Fisher's significance-testing approach. The papers supply much of the standard terminology still used today, including the term "alternative hypothesis" and the H0/H1 notation for the null and alternative hypotheses, and also introduce the simple/composite hypothesis distinction. Fisher and Neyman would go on to quarrel over the relative merits of their competing formulations until Fisher's death in 1962; the two approaches were later merged into a single hybrid framework by textbook writers and practitioners without direct input from either originator.[122] | United Kingdom |
| 1933 | Method | English mathematician Raymond Paley publishes "On orthogonal matrices" in the Journal of Mathematics and Physics, introducing two methods for constructing infinite families of Hadamard matrix using Galois fields — one for orders that are a multiple of 4 where q = N−1 is an odd prime power, and a second for orders N = 2q+2 under related conditions. The two-level orthogonal arrays derived from Paley's Hadamard matrices become known as Paley designs, and his construction underlies the Plackett–Burman designs Robin Plackett and James P. Burman would introduce in 1946.[123] | United Kingdom |
| 1933 | Experiment | At the Royal London Hospital, William Evans and Clifford Hoyle conduct an early comparative trial of drugs used continuously to treat angina pectoris, testing 90 patients by comparing the outcomes of active drug treatment against a dummy ("placebo") treatment within the same trial. Finding no significant difference between the two, the researchers conclude that the drugs tested exert no specific pharmacological effect on the condition — an early instance of the placebo-controlled comparative design that Harry Gold and colleagues would apply at much larger scale four years later.[124] | England |
| 1935 | Method | Jerzy Neyman, with K. Iwaszkiewicz and St. Kolodziejczyk, publishes "Statistical Problems in Agricultural Experimentation" in the Supplement to the Journal of the Royal Statistical Society, extending the randomization-based model Neyman had developed in his 1923 essay "On the Application of Probability Theory to Agricultural Experiments" to cover randomized block designs.[125] | Poland |
| 1935 | Concept | Ronald Fisher introduces the term "confounding" in The Design of Experiments to describe a specific consequence of blocking (statistics) in a factorial experiment — when partitioning treatment combinations into blocks causes certain interaction effects to become indistinguishable from ("confounded with") block effects. Fisher's usage popularizes the term within statistics, though his concern was controlling heterogeneity among experimental units rather than causal inference in the modern sense.[126] | United Kingdom |
| 1935 | Publication | Ronald Fisher publishes The Design of Experiments, a foundational work that formalizes principles such as randomization, replication, and blocking, and emphasizes the importance of efficient experimental design; the book becomes the basis of modern experimental science and remains one of the field's foundational texts.[127][128][129][22][21][19] | United Kingdom |
| 1935 | Experiment | Ronald Fisher describes the "lady tasting tea" experiment in The Design of Experiments, the original exposition of his concept of the null hypothesis — which he describes as never proved or established, only possibly disproved through experimentation. The example is loosely based on a real event: phycologist Muriel Bristol claimed she could tell whether tea or milk was poured into a cup first, and her future husband William Roach suggested testing her with eight cups, four of each preparation, in random order. Fisher's treatment, using what becomes known as Fisher's exact test, calculates that correctly identifying all eight cups would occur by chance alone in only 1 of 70 possible combinations (about 1.4%), below the conventional 5% significance threshold — establishing the logic of randomized assignment combined with a combinatorial significance test as a template later applied throughout experimental science. According to statistician H. Fairfield Smith, as later reported by David Salsburg, Bristol did in fact identify all eight cups correctly in the real trial.[130][131] | United Kingdom |
| 1935 | Concept | The term "factorial" (in the experimental-design sense) appears to enter print for the first time when Ronald Fisher uses it in The Design of Experiments.[132] | United Kingdom |
| 1937 | Method | Researcher Meredith Crawford, at the Yerkes National Primate Research Center, invents the cooperative pulling paradigm — an experimental design in which two or more animals must pull a reward toward themselves via an apparatus neither can operate alone — publishing a study of two young chimpanzees, Bula and Bimba, coordinating to pull ropes attached to a box too heavy for either to move alone. The paradigm becomes the most widely used experimental design for testing cooperation in animals, later applied to bonobos, orangutans, capuchins, elephants, wolves, dogs, ravens, kea, and dolphins, among other species, to investigate the evolution and mechanisms of cooperative behavior.[133] | United States |
| 1937 | Experiment | American cardiologist Harry Gold, with Nathaniel Kwit and Harold Otto, publishes a placebo-controlled trial of xanthine drugs (theobromine and aminophylline) for cardiac pain in the Journal of the American Medical Association, testing roughly 700 subjects — a substantial scale-up from earlier placebo-controlled comparisons such as William Evans and Clifford Hoyle's 1933 angina trial, and part of Gold's broader advocacy (through the Cornell Conferences on Therapy) for masked, placebo-controlled assessment in American clinical pharmacology.[134] | United States |
| 1937 | Method | Frank Yates formalizes what becomes known as the Yates analysis in "The Design and Analysis of Factorial Experiments," a technical communication published by the Commonwealth Bureau of Soils in Harpenden, England. The technique exploits the special structure of full and fractional factorial experiments — arranging response data in a specific "Yates' order" and repeatedly summing and differencing pairs of values — to generate least-squares estimates of factor effects and interactions for all factors simultaneously, without needing to separately fit each model.[135] | United Kingdom |
| 1939 | Concept | A publication by Bose and Nair underlie the concept of association scheme. In their paper, they introduced the concept of association schemes as a way to study the structure of contingency tables. They show that association schemes can be used to represent the dependencies between the variables in a contingency table, and that they can be used to derive statistical tests for independence.[136] | India |
| 1940 | Method | Ronald A. Fisher publishes "An Examination of the Different Possible Solutions of a Problem in Incomplete Blocks" in Annals of Eugenics, extending his earlier work on blocking (statistics) to systematically examine balanced incomplete block designs — designs in which not every treatment can be tested in every block, a situation Fisher's original randomized block design did not address.[137] | United Kingdom |
| 1940 | Method | Raj Chandra Bose and K. Kishen at the Indian Statistical Institute independently find some efficient designs for estimating several main effects. | India |
| 1942 | Concept | K. Kishen generalizes Latin squares and mutually orthogonal Latin squares to Latin cubes and Latin hypercubes in "On latin and hyper-graeco cubes and hypercubes," published in Current Science.[138] | India |
| 1943 | Method | Economist Robert Dorfman introduces group testing in a short report published in the Annals of Mathematical Statistics, motivated by the United States Public Health Service's wartime effort to screen soldiers for syphilis without the expense of testing every individual blood sample separately. Dorfman's method pools blood samples into groups and tests each pooled sample; groups testing negative allow every soldier within them to be cleared with a single test, while positive groups require individual follow-up testing — dramatically reducing the expected number of tests needed when disease prevalence is low. Dorfman tabulates the optimal group size as a function of prevalence rate.[139] | United States |
| 1944 | Experiment | The Patulin Clinical Trials Committee of the Medical Research Council (United Kingdom) conducts a multicentre trial of the antibiotic patulin on the course of the common cold. Long overlooked, the trial has more recently been identified by historians of medicine as the first properly controlled, randomized multicentre trial conducted under the Council's aegis — predating the Medical Research Council (United Kingdom)'s better-known 1948 streptomycin tuberculosis trial, which had long been popularly credited as the first randomized clinical trial.[140][141] | United Kingdom |
| 1945 | Method | American statistician Abraham Wald publishes Sequential Tests of Statistical Hypotheses, pioneering the field of sequential analysis — a framework for statistical tests where the number of observations is not fixed in advance, and the decision to stop collecting data (accept, reject, or continue sampling) can depend on the accumulated results as the experiment proceeds. Wald's work, developed partly in a wartime context, provides the formal apparatus later applied throughout industrial and clinical sequential experimental design.[142] | United States |
| 1945 | Concept | British statistician D. J. Finney publishes "The Fractional Replication of Factorial Arrangements" in Annals of Eugenics, containing the first statistical use in print of the term "aliasing (factorial experiments)" — the phenomenon in fractional factorial designs where some effects become indistinguishable from each other because only a fraction of all possible treatment combinations is observed. The term would later be adopted into signal processing theory, possibly influenced by this earlier statistical usage.[143] | United Kingdom |
| 1945 | Method | British statistician D. J. Finney introduces fractional factorial design, extending Ronald Fisher's earlier work on full factorial experiments at the Rothamsted Experimental Station by showing how to test only a carefully chosen fraction of all possible factor-level combinations — reducing the number of experimental runs required while deliberately confounding selected effects, based on the assumption that higher-order interactions are typically negligible (the sparsity-of-effects principle). Originally developed for agricultural applications, fractional factorial design later spreads to engineering, business, and other sciences.[144] | United Kingdom |
| 1946 | Experiment | Psychologist E.M. Jellinek conducts a crossover trial for a U.S. headache-remedy manufacturer testing whether removing a scarce wartime ingredient would reduce drug efficacy, randomly assigning 199 subjects with frequent headaches to rotate through four formulations (three real combinations of ingredients plus a lactate placebo) over eight weeks. Initial analysis across all subjects suggests the scarce ingredient is unnecessary, but Jellinek discovers that 120 of the 199 subjects are "placebo reactors" whose inclusion masks the ingredient's real effect; restricting analysis to the 79 non-reactors reveals the ingredient does contribute significantly to efficacy. The trial becomes an influential early demonstration that placebo responsiveness varies systematically across individuals and can confound drug-efficacy comparisons if not accounted for.[145] | United States |
| 1946 | Policy | The Cornell Conferences on Therapy, a series of consensus meetings among American clinical researchers, decree that new pharmaceutical substances should be tested against placebos under blinded, randomized conditions — an early institutional codification, ahead of any government regulation, of what would become the standard modern clinical trial design.[119] | United States |
| 1946 | Method | R.L. Plackett and J.P. Burman publish a renowned paper titled "The Design of Optimal Multifactorial Experiments". The paper introduces what would be called Plackett–Burman designs, which are highly efficient screening designs with run numbers that are multiples of four. These designs are particularly useful for experiments where only main effects are of interest. In a Plackett-Burman design, main effects are often heavily confounded with two-factor interactions, making them suitable for screening experiments. For instance, a Plackett-Burman design with 12 runs can be utilized for an experiment containing up to 11 factors.[146] | United Kingdom |
| 1946–1947 | Concept | C. R. Rao generalizes Kishen's hypercube constructions to arrays of arbitrary strength t in "Hypercubes of strength d leading to confounded designs in factorial experiments" (1946), then introduces the general concept of the orthogonal array — a tabular structure in which every selection of t columns contains all possible t-tuples of symbols the same number of times — in "Factorial experiments derivable from combinatorial arrangements of arrays" (1947), the paper credited as the origin of the modern notion of orthogonal arrays used throughout combinatorial design theory, coding theory, cryptography, and software testing.[147][148] | India |
| 1946–1948 | Experiment | Sir Geoffrey Marshall of the Medical Research Council's Tuberculosis Research Unit leads the first randomized curative clinical trial, testing the efficacy of streptomycin for pulmonary tuberculosis under double-blind, placebo-controlled conditions — though the trial is more commonly credited to the Medical Research Council (United Kingdom) as an institution than to its individual lead investigator. Results, published in 1948, show reduced mortality alongside emerging drug resistance and side effects, establishing modern clinical trial methods such as random allocation and blinding and shaping the future of evidence-based medicine.[149][150] | United Kingdom |
| 1947 | Policy | The Nuremberg Code is articulated by the U.S. military tribunal in USA v. Brandt (the "Doctors' Trial"), part of the Subsequent Nuremberg Trials against German physicians who conducted unethical experiments on concentration-camp prisoners and carried out over 3.5 million forced sterilizations. Concerned that the defendants — who argued their experiments differed little from prewar research and that no law distinguished legal from illegal experimentation — might escape conviction, prosecution medical experts Andrew Conway Ivy and Leo Alexander drafted a memorandum in April 1947 outlining principles for legitimate medical research; an expanded version requiring explicit voluntary consent followed on 9 August 1947. The judges' verdict, delivered 20 August 1947 against Karl Brandt (physician) and 22 others, revised these into ten points establishing principles including informed consent, freedom from coercion, and beneficence toward research subjects — the first codification of research ethics standards, influencing medical experiment codes of practice worldwide despite going largely unenforced (and even dismissed by some as inapplicable to "ordinary physicians") for roughly two decades after it was written.[151][152] | Germany |
| 1948 | Concept | Francis J. Anscombe, at Rothamsted Experimental Station, discusses and develops design-based (randomization-based) analysis of variance, later extended by Oscar Kempthorne at Iowa State University, who introduces the assumption of unit-treatment additivity — the idea that an experimental unit's observed response can be written as the sum of the unit's baseline response and a treatment effect that is constant across all units receiving that treatment. This randomization-based approach differs from the more commonly taught normal-linear-model approach in making no assumption of a normal distribution or of independence between observations, though the two approaches' test statistics closely approximate each other in practice.[153] | United Kingdom |
| 1948 | Concept | British statistician Frank Yates introduces the concept of restricted randomization.[154][155] | United Kingdom |
| 1949 | Method | Psychologist Richard Solomon (psychologist) develops the Solomon four-group design, a research method addressing the problem of pretest sensitization — the possibility that administering a pre-intervention test itself influences how subjects respond to a subsequent treatment. In addition to the standard pre-test/treatment/post-test and pre-test/control/post-test groups, the design adds two further groups that skip the pre-test entirely (treatment-only and control-only, each still post-tested), allowing researchers to separate the effect of the treatment itself from the effect of having taken the pre-test.[156] | United States |
| 1949 | Method | Genichi Taguchi begins developing his experimental design techniques while working at Japan's Electrical Communications Laboratories (ECL) in the post–World War II reconstruction period. Tasked with improving research and development productivity, he formulates a new approach to quality improvement that emphasizes off-line quality control, robust design, and statistical experimentation. These early efforts lay the foundations of what would later become known as the Taguchi Methods, integrating engineering design with statistical principles to systematically reduce variability, improve product quality, and lower societal and manufacturing costs.[157] | Japan |
| 1949 | Method | Kenneth Arrow, David Blackwell, and M.A. Girshick publish "Bayes and Minimax Solutions of Sequential Decision Problems" in Econometrica, an early and influential contribution to sequential decision theory building on Abraham Wald's wartime work on sequential analysis, formalizing Bayes and minimax approaches to problems where the decision of when to stop sampling is itself part of the optimization.[158] | United States |
| 1949 | Concept | R. H. Bruck and H. J. Ryser prove a nonexistence result for finite projective planes — special cases of symmetric block designs with λ = 1 — showing that if a projective plane of order q exists and q ≡ 1 or 2 (mod 4), then q must be expressible as the sum of two squares. The result rules out projective planes of certain orders (such as 6) that satisfy the basic combinatorial parameter equations for a symmetric design but are nonetheless impossible.[159] | United States |
| 1950 | Publication | Gertrude Mary Cox and William Gemmell Cochran publish Experimental Designs, following the tradition established by Ronald Fisher's 1935 The Design of Experiments as a practical reference for statisticians and scientists. The book opens with three chapters on the philosophy of statistics in experimentation before covering specific designs — completely randomized designs, randomized blocks, Latin squares, factorial designs, split-plot designs, and various incomplete block designs — illustrated throughout with results from actual experiments. Cox, a pioneering woman in a field with little female visibility at the time, had helped found the Statistics Laboratory at Iowa State and was the first head of the Department of Experimental Statistics at North Carolina State; Cochran had worked at Rothamsted under Frank Yates before moving to Iowa State. Oscar Kempthorne would soon follow with a complementary theoretical treatment, The Design and Analysis of Experiments (1952), and Cochran and Cox would publish a second edition of their own book in 1957.[160] | United States |
| 1950 | Method | P.M. Grundy and Michael Healy (statistician) publish "Restricted Randomization and Quasi-Latin Squares" in the Journal of the Royal Statistical Society, Series B, developing restricted randomization — a method for excluding intuitively poor or undesirable treatment allocations from the space of possible randomizations (for example, preventing a new obesity treatment from being randomly assigned only to the heaviest patients) while still preserving the theoretical statistical benefits of randomization. The concept had also been introduced independently by Frank Yates, whose earlier work on the subject this timeline's 1948 entry describes.[161] | United Kingdom |
| 1950 | Concept | K. Bush coins the term "orthogonal array" in his PhD thesis at the University of North Carolina, naming the structure C. R. Rao had introduced in 1947 (Rao himself had used the unmodified term "array," meaning simply a subset of treatment combinations, before the tabular/matrix framing made the possibility of non-simple arrays — those with repeated rows — apparent).[162] | United States |
| 1950 | Method | Cochran and Cox formalize principles of sample size determination in experimental design, providing more accurate tables and methods.[19] | United States |
| 1950 | Experiment | Richard Doll and Austin Bradford Hill publish a preliminary report demonstrating a statistically significant association between tobacco smoking and lung cancer, using a large case–control study design comparing lung cancer patients against controls without the disease. Critics initially argued the case–control design could not establish causation; subsequent cohort studies by the same researchers over following decades confirmed the causal link the case–control study had suggested, and smoking is now recognized as the cause of roughly 87% of lung cancer deaths in the United States.[163] | United Kingdom |
| 1950 | Concept | S. Chowla and H. J. Ryser extend the 1949 Bruck–Ryser nonexistence result from projective planes to general symmetric block designs, establishing what becomes known as the Bruck–Ryser–Chowla theorem: necessary conditions (a square-difference condition when the number of points is even, and a solvable Diophantine equation when odd) that any symmetric (v, b, r, k, λ)-design's parameters must satisfy, ruling out entire classes of otherwise combinatorially plausible designs.[164] | United States |
| 1950 | Concept | Randomized controlled trials begin to emerge as the gold standard in medical research, enabling systematic causal inference | Global |
| 1951 | Method | Formal placebo-controlled and multi-arm clinical trial designs are developed to improve reliability of treatment comparisons | United States |
| 1951 | Method | George E. P. Box and K. B. Wilson introduce response surface methodology (RSM), using sequences of designed experiments and second-degree polynomial models to approximate and optimize responses, even with limited process knowledge. They develop a method to build a quadratic model where the number of data points scales linearly, rather than exponentially, with the number of inputs, striking a balance between accuracy and ease of application — enabling sequential, adaptive optimization of industrial processes even under limited prior knowledge of the underlying system.[165][166][14][22][20] | United Kingdom |
| 1952 | Concept | American mathematician and statistician Herbert Robbins recognizes the significance of a problem where a gambler faces a trade-off between "exploitation" of the machine with the highest expected payoff and "exploration" to learn about other machines' payoffs. This problem involves pulling levers on different machines, each providing random rewards from unknown probability distributions. The gambler aims to maximize the total rewards earned over a sequence of lever pulls. Robbins devised convergent population selection strategies in his work on "some aspects of the sequential design of experiments."[167] | United States |
| 1952 | Method | D.G. Horvitz and D.J. Thompson publish "A Generalization of Sampling Without Replacement from a Finite Universe" in the Journal of the American Statistical Association, introducing what becomes known as the Horvitz–Thompson estimator — a method for producing unbiased population estimates from samples with unequal selection probabilities by weighting each observation by the inverse of its probability of inclusion. Originally developed for survey sampling, the estimator later becomes a standard tool via inverse probability weighting for correcting bias in the analysis of experiments with unequal exposure probabilities, including studies of spillover and network interference effects.[168] | United States |
| 1952 | Concept development | Bose and Shimamoto introduce the term association scheme.[169] | United States |
| 1954 | Experiment | The Salk polio vaccine trial becomes one of the first large-scale randomized controlled trials, involving over one million participants | United States |
| 1954 | Concept | Paul Lazarsfeld and others formalize factor analysis as a tool for latent variable modeling in experiments. | United States |
| 1954 | Publication | Edwin Boring discusses the history and meanings of control treatments, clarifying the role of controls in experimental research.[19] | United States |
| 1954 | Publication | American experimental psychologist Edward Boring writes an article titled The History of Experimental Design. In this article, Boring notes that the early history of ideas on the planning of experiments has been "but little studied".[101] | United States |
| 1955 | Experiment | An influential study entitled The Powerful Placebo firmly establishes the idea that placebo effects are clinically important.[170] | United States |
| 1955 | Method | M. B. Wilk introduces the randomization analysis of the generalized randomized block design (GRBD), which replicates each treatment at least twice within each block — unlike the classic randomized block design, which has no within-block replication — allowing block-treatment interaction (statistics) to be estimated and tested without relying on parametric assumptions about the distribution of experimental error.[171] | United States |
| 1956 | Concept | D. V. Lindley publishes "On a Measure of Information Provided by an Experiment," proposing that the utility of an experimental design be measured as the Kullback–Leibler divergence between the posterior and prior probability distributions over the parameters being estimated — showing this expected utility is coordinate-independent and equals the mutual information between the parameters and the observations, exactly the expected information gain from running the experiment. This becomes a foundational formulation within Bayesian experimental design.[172] | United Kingdom |
| 1956 | Experiment | Richard Doll and Austin Bradford Hill publish a second report on the mortality of British doctors, a prospective cohort study confirming the causal link between smoking and lung cancer that their 1950 case–control study had first suggested — an early demonstration of how case–control findings could be validated through a stronger, prospective study design.[173] | United Kingdom |
| 1959 | Concept | Leslie Kish uses the term "confounding" in the sense of "incomparability" between two or more groups (such as exposed and unexposed groups) in an observational study — a distinct usage from Fisher's blocking-based sense, and closer to the modern causal-inference meaning of the term as it would later develop in epidemiology.[174] | United States |
| 1959 | Method | Raymond H. Pierson and Edward A. Fay publish "Guidelines for Interlaboratory Testing Programs" in Analytical Chemistry, an early formalization of round-robin testing — interlaboratory studies in which multiple independent laboratories perform the same measurement or analysis, using the same or different methods and equipment, to assess the reproducibility of a test method or verify a new method of analysis against an established one.[175] | United States |
| 1959 | Concept | Raj Chandra Bose and D. M. Mesner publish "On Linear Associative Algebras Corresponding to Association Schemes of Partially Balanced Designs" in the Annals of Mathematical Statistics, introducing what becomes known as the Bose–Mesner algebra — the commutative algebra generated by the adjacency matrices of an association scheme's relations. This publication turns association schemes, previously a statistical tool for classifying partially balanced incomplete block designs, into an object of independent algebraic interest.[176] | United States |
| 1959 | Concept | W. M. S. Russell and R. L. Burch publish The Principles of Humane Experimental Technique, introducing the "Three Rs" — Replacement (preferring non-animal methods when they can achieve the same scientific aims), Reduction (obtaining comparable information from fewer animals, or more information from the same number), and Refinement (minimizing pain, suffering, or distress and enhancing welfare for animals that must be used) — as guiding principles for more ethical animal research. The Three Rs would go on to become explicit in animal-research legislation in many countries and remain foundational to the ethical design of experiments involving animals.[177] | United Kingdom |
| 1959 | Concept | Arthur Samuel publishes "Some Studies in Machine Learning Using the Game of Checkers" in IBM Journal of Research and Development, describing a self-improving checkers program that, given only the rules of the game and a list of parameters of unknown relative importance, learns within 8–10 hours of play to outperform the person who wrote it. The paper is widely credited with popularizing the term "machine learning," though the commonly quoted definition attributed to Samuel — giving computers "the ability to learn without being explicitly programmed" — is now generally considered apocryphal, as it does not appear verbatim in his published work.[178] | United States |
| 1959–1961 | Method | Jack Kiefer and Jacob Wolfowitz develop formal optimal design theory, introducing criteria-based selection of experimental designs for maximum precision.[20] | United States |
| 1960 | Industrial statistics | Japanese engineer and statistician Genichi Taguchi publishes Design of Experiments for Engineers, advancing statistical methods for industrial experimentation. His work helps formalize what later would become known as Taguchi methods, integrating experimental design and statistical analysis to improve product quality, optimize manufacturing processes, and reduce costs in engineering and industry.[179] | Japan |
| 1960 | Method | George E. P. Box and Donald Behnken publish "Some New Three Level Designs for the Study of Quantitative Variables" in Technometrics, devising what becomes known as Box–Behnken designs — three-level experimental designs for response surface methodology that place each factor at one of three equally spaced values and are efficient enough to fit a quadratic model. The seven-factor design, whose estimation variance depends almost exactly on distance from the centre point, was found first; designs for other numbers of factors were subsequently derived to approximate this same rotatability property.[180][181] | United States |
| 1960 | Method | Damaraju Raghavarao publishes "Some Aspects of Weighing Designs" in the Annals of Mathematical Statistics, establishing the statistical theory connecting weighing matrices to the classical problem of estimating the weights of multiple objects with minimum variance using a balance scale. Raghavarao shows that the variance of least-squares weight estimates is minimized if and only if the design matrix used to allocate objects to the scale's pans is a proper weighing matrix, connecting a combinatorial mathematics object to a concrete experimental design problem.[182] | United States |
| 1960 | Experiment | English psychologist Peter Wason conducts an experiment in which participants must discover a rule governing triples of numbers, given that (2,4,6) fits it. Participants overwhelmingly test only triples that confirm their current hypothesis rather than triples designed to falsify it, and rarely discover the true (very broad) rule, "any ascending sequence." Wason interprets the results as showing a preference for confirmation over falsification in hypothesis testing, coining the term "confirmation bias."[183] | United Kingdom |
| 1960 | Method | Donald Thistlethwaite and Donald T. Campbell introduce the regression discontinuity design in "Regression-Discontinuity Analysis: An alternative to the ex post facto experiment," published in the Journal of Educational Psychology, applying it to evaluate the effect of merit-based scholarship programs on students' career plans. The design estimates a local causal treatment effect by comparing outcomes for units lying just above and just below a predetermined threshold that determines treatment assignment (e.g. a scholarship awarded only to students scoring above a cutoff), exploiting the idea that units on either side of a narrow cutoff should otherwise be comparable even without formal random assignment. It becomes a widely used quasi-experimental design across economics, political science, epidemiology, and psychology.[184] | United States |
| 1961 | Concept | George E. P. Box and J. S. Hunter introduce the concept of resolution in their paper "The 2^(k-p) Fractional Factorial Designs," providing a way to measure the degree to which a fractional factorial design avoids aliasing between main effects and important interactions — a design has resolution R if every effect involving p factors is unaliased with every effect having fewer than R−p factors. This concept becomes a standard way of summarizing and comparing the aliasing properties of fractional designs.[185] | United States |
| 1961 | Experiment | Ecologist Joseph H. Connell publishes "The Influence of Interspecific Competition and Other Factors on the Distribution of the Barnacle Chthamalus stellatus" in Ecology, a field study examining how competition from the barnacle species Balanus balanoides shapes the distribution of Chthamalus stellatus on rocky shores. The study becomes a widely cited early example of manipulative field experimentation in ecology.[186] | England |
| 1961 | Concept | The term nocebo (Latin nocēbō, "I shall harm", from noceō, "I harm")[187] is coined by Walter Kennedy to denote the counterpart to the use of placebo (Latin placēbō, "I shall please", from placeō, "I please"; a substance that may produce a beneficial, healthful, pleasant, or desirable effect). Kennedy emphasized that his use of the term "nocebo" refers strictly to a subject-centered response, a quality inherent in the patient rather than in the remedy".[188] | United States |
| 1962 | Policy | The Kefauver Harris Amendment requires proof of efficacy through controlled trials before drug approval, institutionalizing experimental design in regulation. | United States |
| 1962 | Experiment | Researchers including Walter N. Pahnke conduct the Marsh Chapel Experiment at Boston University, a double-blind study giving divinity students either the psychedelic substance psilocybin or an active placebo — a large dose of niacin chosen specifically because it produces noticeable physical sensations that could lead control subjects to believe they too had received a psychoactive drug. The design becomes an influential early example of using an active rather than inert placebo to counter unblinding in trials of drugs with strong subjective effects.[189] | United States |
| 1962 | Method | Japanese statistician Kazumasa Kôno publishes "Optimum designs for quadratic regression on k-cube" in the Memoirs of the Faculty of Science, Kyushu University, deriving optimal experimental designs for quadratic response surface methodology models that require fewer experimental runs than George E. P. Box's earlier central-composite designs. Together with later work by Jack Kiefer (mathematician), the Kôno–Kiefer analysis explains why such optimal designs, despite their mathematical construction as probability measures over continuous spaces, can be supported on a small number of discrete points closely resembling the traditional designs already used in practice.[190] | Japan |
| 1962 | Experiment | Vernon L. Smith publishes An Experimental Study of Competitive Market Behavior in the Journal of Political Economy. Using controlled laboratory market experiments with human participants, he tests hypotheses of neoclassical competitive theory and demonstrates price convergence toward equilibrium, helping establish experimental economics as a systematic empirical research methodology.[191] | United States |
| 1962 | Method | British statistician John Nelder proposes a set of systematic, circular experimental designs as an alternative to the replicated, full factorial spacing experiments. These designs, known as the Nelder 'wheel' design, are developed to address limitations related to space and plant material. The design consists of a circular plot with concentric circumferences radiating outward, connected by spokes that extend from the center to the farthest circumference. Trees are planted at the intersections of spokes and circumferences within the plot.[192] | United Kingdom |
| 1963 | Concept | Campbell and Stanley discuss design according to the categories of preexperimental designs, experimental designs, and quasi-experimental designs.[193] | United States |
| 1963 | Method | J.J. Boren introduces "repeated acquisition of new behavioral chains" as a single-subject research design in the American Psychologist, offering an alternative to reversal designs for studying behaviors that cannot practically or ethically be reversed once learned — instead repeatedly measuring how quickly a subject acquires a new behavioral sequence under different experimental conditions.[194] | United States |
| 1964 | Policy | The World Medical Association adopts the Declaration of Helsinki at its 18th World Medical Assembly in Helsinki, Finland, setting out "Recommendations Guiding Physicians in Biomedical Research Involving Human Subjects." The Declaration requires that research protocols be reviewed by an independent committee, that subjects give freely obtained informed consent after being told the study's aims, methods, and risks, and that the subject's wellbeing always take precedence over the interests of science and society. It becomes a foundational reference document for research ethics committees worldwide and would be amended repeatedly in subsequent decades, including in Tokyo (1975) and Venice (1983).[195] | Finland |
| 1965 | Concept | Leslie Kish coins the term "design effect" in his book Survey Sampling, proposing the general definition as the ratio of the variance of an estimator under a given (often complex) sampling design to its variance under simple random sampling of the same size — along with formulas for the design effect under cluster sampling (incorporating intraclass correlation) and under unequal-probability sampling. These become known as "Kish's design effect" and remain foundational to survey methodology.[196] | United States |
| 1966 | Concept | Hanan Selvin and Alan Stuart describe "data-dredging procedures" in survey analysis, formalizing an early statistical critique of selecting which explanatory variables to retain in a model based on the same data used to test them. Using the metaphor of fish that "don't fall through the net," they argue that variables retained after such a selection process are systematically biased toward appearing more significant than they truly are, altering the validity of standard statistical tests applied afterward.[197] | United Kingdom |
| 1966 | Publication | American psychologist Robert Rosenthal (psychologist) publishes Experimenter Effects in Behavioral Research, a foundational study establishing that researchers can subtly and unconsciously communicate their expectations to experimental subjects, biasing outcomes toward those expectations — later demonstrated experimentally by showing that researchers told to expect positive ratings from subjects obtained significantly more positive data than researchers told to expect negative ratings, using the identical task.[198] | United States |
| 1971 | Publication | Raymond H. Myers publishes Response Surface Methodology, a textbook establishing widely used practical guidelines for central composite design parameters — including standard methods for selecting the axial-point distance α (orthogonal and rotatable design criteria) — that remain in use in statistical software decades later.[199] | United States |
| 1972 | Publication | Herman Chernoff writes an overview of optimal sequential designs[200] In the design of experiments, optimal designs is a class of experimental designs that are optimal with respect to some statistical criterion. | United States |
| 1973 | Concept | Belgian mathematician Philippe Delsarte's thesis, "An Algebraic Approach to the Association Schemes of Coding Theory," recognizes and develops the deep connections between association schemes and both coding theory and design theory, becoming widely regarded as the most important contribution to association scheme theory after its statistical origins in design of experiments.[201] | Belgium |
| 1973 | Method | Donald Rubin publishes "Matching to Remove Bias in Observational Studies" in Biometrics, formalizing statistical matching — pairing treated units in an observational study with non-treated units sharing similar observable characteristics — as a technique for reducing confounding bias in estimated treatment effects when random assignment is unavailable. The paper predates and lays groundwork for Rubin's later development of propensity score matching with Paul Rosenbaum in 1983.[202] | United States |
| 1973 | Concept | Psychologist Edwin Wike publishes "Water beds and sexual satisfaction: Wike's law of low odd primes (WLLOP)" in Psychological Reports, a humorous methodological note observing that whenever an experiment has a low odd prime number of treatment conditions (three being the most common case), the design tends to be unbalanced and partially confounded. Wike illustrates the principle with an invented three-group study of water beds and sexual satisfaction, showing that a fourth group is needed to disentangle the effect of the water bed from that of an accompanying seasickness pill. Wike explicitly notes that naming the "law" after himself is itself an example of Stigler's law of eponymy.[203] | United States |
| 1974 | Concept | Donald Rubin formalizes the potential outcomes framework, defining causal effects using counterfactual comparisons between treated and untreated states for the same unit.[204] | United States |
| 1974 | Method | R.H.F. Denniston solves a problem posed by James Joseph Sylvester in 1860 as an extension of Kirkman's schoolgirl problem: whether thirteen disjoint Steiner triple systems S(2,3,15) exist, such that "Kirkman's schoolgirls" could march in triples for an entire 13-week term without any pair of girls being grouped together twice. Denniston's computer search, run for seven hours on an Elliott 4130 computer at the University of Leicester, finds a valid week-one solution from which all subsequent weeks are generated by a fixed relabeling scheme; the number of non-isomorphic solutions to Sylvester's problem remains unknown.[205] | United Kingdom |
| 1975 | Concept | Dijen K. Ray-Chaudhuri and R. M. Wilson prove a generalization of Fisher's inequality to t-designs, showing that in a 2s-(v, k, λ) design the number of blocks is at least the binomial coefficient v choose s — extending the classical Fisher's inequality (b ≥ v for balanced incomplete block designs) to a broader class of combinatorial designs.[206] | United States |
| 1975 | Concept | A.C. Atkinson and V.V. Fedorov publish "The design of experiments for discriminating between two rival models" in Biometrika, introducing T-optimality — a criterion for constructing experimental designs specifically intended to maximize the statistical discrepancy between two competing candidate models at the chosen design points, rather than to minimize estimation variance within a single assumed model. The criterion becomes particularly important in biostatistics applications supporting pharmacokinetics and pharmacodynamics, building on earlier work by David R. Cox and Atkinson on model-discrimination experiments.[207] | United Kingdom |
| 1975 | Method | Stuart Pocock and Richard Simon introduce minimisation in Biometrics, an adaptive stratified sampling technique for balancing treatment groups across multiple prognostic factors in clinical trials. Unlike traditional blocking (statistics), which requires a separate randomisation list for each combination of stratification factors — a number of lists that grows exponentially as more factors are added — minimisation calculates, for each new patient, the imbalance that would result from allocating them to each treatment group, sums the imbalance across all factors, and assigns the patient to whichever group minimises the overall imbalance (optionally with a random element retained). The method comes to be described by some as maintaining better balance than blocked randomisation, particularly as the number of stratification factors grows.[208] | United Kingdom |
| 1976 | Method | James Scheirer, William S. Ray, and Nathan Hare publish "The Analysis of Ranked Data Derived from Completely Randomized Factorial Designs" in Biometrics, introducing the Scheirer–Ray–Hare test, a non-parametric extension of the Kruskal–Wallis one-way analysis of variance for examining whether a measured outcome is affected by two or more factors, without requiring the assumption of a normal data distribution. Though more conservative in statistical power than parametric multi-factorial ANOVA and later described than most comparable variance-analysis tests, the Scheirer–Ray–Hare test finds continued use, particularly in the biological sciences, as a non-parametric alternative when normality assumptions cannot be met.[209] | United States |
| 1976 | Publication | Douglas C. Montgomery publishes Design and Analysis of Experiments, a comprehensive textbook on the design and analysis of experiments. The book covers a wide range of topics, including principles of experimental design, different types of experimental designs, analysis of experimental data, and use of experimental design in a variety of fields, such as agriculture, industry, and medicine.[210] | United States |
| 1976 | Method | Neil Sloane and Martin Harwit publish "Masks for Hadamard transform optics, and weighing designs" in Applied Optics, extending weighing-matrix theory from its origin in physical balance-scale weighing to the design of optical masks used in spectrometers and image scanners — where each matrix element determines whether incoming light is transmitted to one detector, absorbed, or reflected to a second detector, mathematically equivalent to the balance-pan weighing problem.[211] | United States |
| 1976 | Method | Olli Miettinen formalizes the conditions under which the odds ratio of exposure in a case–control study can be used to estimate relative risk, refining earlier work by Jerome Cornfield that had shown this approximation held specifically when the disease outcome under study is rare.[212] | United States |
| 1977 | Method | Stuart Pocock introduces group sequential methods for the design and analysis of clinical trials, in which patient entry is divided into equal-sized groups so that repeated significance tests on the accumulated data after each group determine whether the trial continues; the associated stopping threshold becomes known as the Pocock boundary.[213] | United Kingdom |
| 1977 | Method | Latvian engineer Vilnis Eglājs and P. Audze propose a space-filling multifactor experimental design technique — later known as the Audze–Eglājs (AE) method — in Russian-language literature, minimizing a potential-energy-like criterion between design points to spread them evenly across the design space. The technique is later recognized as equivalent in spirit to Latin hypercube sampling, predating Michael McKay's independently-derived and far more widely cited 1979 description of Latin hypercube sampling by two years, though Eglājs and Audze's work remains comparatively obscure in the Western statistical literature, likely due to the Russian-language original and limited circulation outside Soviet-bloc engineering circles.[214][215] | Latvia |
| 1978 | Concept | Donald Rubin introduces the concept of "ignorable assignment mechanisms" in causal inference, formalizing the condition under which the way individuals were assigned to treatment groups can be disregarded during statistical analysis, given everything else recorded about them — a refinement of his earlier Rubin causal model potential-outcomes framework, and part of what becomes known as the Neyman-Rubin causal inference model.[216] | United States |
| 1978 | Concept | According to Box et al., experimental design refers to the systematic layout of combinations of variables. The layouts in the case of concepts are test concepts or test vignettes.[217] | United States |
| 1978 | Concept | Ulrich Krengel and Louis Sucheston (with David J. H. Garling) formulated the Prophet Inequality in optimal stopping theory. They show that a gambler observing sequential random rewards can secure at least half the expected payoff of a “prophet” who knows all outcomes in advance, establishing a foundational result in probability theory and decision processes.[218] | United States |
| 1979 | Method | Marvin Zelen publishes his new method, which would later be called Zelen's design.[219][220] | United States |
| 1979 | Method | Harvard biostatistician Marvin Zelen (biostatistician) proposes a new randomized clinical trial design in "A New Design for Randomized Clinical Trials," published in the New England Journal of Medicine, in which patients are randomized to treatment or control before informed consent is sought — allowing consent to be obtained conditionally, and control-group patients to be enrolled with minimal disruption to standard care. The design is intended to reduce clinician discomfort with randomization and patient-side effects such as resentful demoralization and the Hawthorne effect, but draws criticism for its lack of allocation concealment and, in trials on serious or life-threatening conditions, for being seen by participants as "inappropriately deceptive and manipulative."[221] | United States |
| 1979 | Publication | Thomas D. Cook and Donald T. Campbell publish Quasi-experimentation: Design & Analysis Issues for Field Settings, extending Campbell and Julian C. Stanley's 1963 categorization of experimental, quasi-experimental, and pre-experimental designs into a comprehensive treatment of quasi-experimental methodology for applied field research where randomization is impractical or unethical. The book becomes a foundational reference for quasi-experimental design across social science, public health, education, and policy analysis, and is substantially revised and expanded by Cook, Campbell, and William R. Shadish in 2002.[222] | United States |
| 1979 | Method | Peter C. O'Brien and Thomas R. Fleming publish "A Multiple Testing Procedure for Clinical Trials" in Biometrics, introducing the O'Brien–Fleming boundary for group sequential design clinical trial monitoring. The boundary sets a very conservative (high) significance threshold at early interim analyses, becoming progressively less stringent as the trial proceeds until it approaches the nominal significance level (e.g. 0.05) at the final analysis — protecting against premature stopping due to random early fluctuations while still preserving the overall Type I error rate specified for the trial. Despite requiring the number and timing of interim analyses to be prespecified, the O'Brien–Fleming boundary becomes one of the most widely used methods for monitoring clinical trials.[223] | United States |
| 1979 | Method | Michael McKay at Los Alamos National Laboratory makes a significant contribution to the field of statistical sampling by introducing the concept of latin hypercube sampling.[224] | United States |
| 1979 | Concept | John C. Gittins publishes "Bandit Processes and Dynamic Allocation Indices," proving that the optimal solution to the multi-armed bandit problem — first posed as an unsolved question in Herbert Robbins's 1952 paper on sequential experimental design — takes the form of an index policy: at each stage, choosing the option (or "arm") with the highest computable "dynamic allocation index," now known as the Gittins index. The result resolves optimal-stopping questions in clinical trial design that had been open since the 1940s, and independent economist Martin Weitzman would establish the equivalent result in economics the same year.[225] | United Kingdom |
| 1980 | Concept | Researchers publish long-term results of the Coronary Drug Project, a study of drugs for long-term treatment of coronary heart disease in men, reporting that participants in the placebo group who adhered to their placebo regimen as instructed showed nearly half the mortality rate of those who did not adhere — despite the placebo itself being pharmacologically inert. The finding becomes a widely cited illustration of the "healthy adherer" effect, in which apparent treatment benefits attributed to adherence may instead reflect underlying differences between compliant and non-compliant patients (health consciousness, psychological effects of following a protocol, or general diligence), a confound later replicated in women with nearly 2.5 times greater survival among placebo-adherent patients.[226] | United States |
| 1980 | Concept | Statistician John Tukey writes about the choice between confirmatory analysis (testing or rejecting existing hypotheses) and exploratory analysis (searching for new hypotheses) in statistical practice, examining how practicing statisticians decide between the two modes of reasoning at different stages of research — an influential early articulation, within statistics specifically, of the same exploratory/confirmatory distinction later applied more broadly to human reasoning by psychologists Jennifer Lerner and Philip Tetlock in 2002.[227] | United States |
| 1981 | Concept | Allen Neuringer first proposes the idea of using single case designs (sometimes referred to as n-of-1 trials) for self-experimentation.[228] | United States |
| 1982 | Literature | British statistician George Box publishes Improving Almost Anything: Ideas and Essays, which gives many examples of the benefits of factorial experiments.[229] | United States |
| 1983 | Concept | Donald Rubin and Paul R. Rosenbaum define the stronger condition of a treatment assignment being "strongly ignorable" in their paper on propensity score matching, establishing propensity scores — the probability of treatment assignment given observed covariates — as central to estimating causal effects from observational data when random assignment is unavailable.[230] | United States |
| 1983 | Method | K.K. Gordon Lan and David L. DeMets publish "Discrete Sequential Boundaries for Clinical Trials" in Biometrika, proposing a method that allows the boundary values of a group sequential trial to be allocated dynamically as the study progresses, addressing a key limitation of the O'Brien–Fleming boundary and other fixed methods that require prespecifying the number of interim analyses and the proportion of total information used at each one.[231] | United States |
| 1984 | Method | French pharmacologist Bernard Bégaud describes the challenge–dechallenge–rechallenge protocol as one of the standardized methods used in France for assessing adverse drug reactions, monitoring whether an adverse event resolves on withdrawal of a medication (dechallenge) and recurs on its re-administration (rechallenge) — a design suited to idiosyncratic, individual-level reactions where conventional population-level statistical testing is unsuitable.[232] | France |
| 1984 | Concept | Stuart Hurlbert publishes a paper in Ecological Monographs where he analyzes 176 experimental studies in ecology. He discovers that 27% of these studies suffer from 'pseudoreplication,' meaning they use statistical testing in situations where treatments are not replicated or replicates were not independent. When considering only studies that use inferential statistics, the percentage of pseudoreplication increases to 48%. To address this issue, Hurlbert suggests interspersing treatments in experiments, even if it means sacrificing randomized samples, particularly in smaller experiments. This approach aims to overcome the problem of pseudoreplication in ecological studies.[233] | United States |
| 1985 | Experiment | Belgian experimental psychologist Jozef Nuttin publishes "Narcissism beyond Gestalt and awareness: the name letter effect" in the European Journal of Social Psychology, using a yoked control design in which pairs of subjects each rate the same sequence of individual letters for preference, unaware that some letters are drawn from their own name and others from their partner's. Because both members of a yoked pair see an identical stimulus sequence, any systematic difference in preference between them can only be attributed to whether a given letter appears in their own name. Nuttin finds that people reliably prefer the letters in their own name over other letters, without being aware of doing so — establishing what becomes known as the name-letter effect.[234] | Belgium |
| 1986 | Experiment | Robert LaLonde finds that findings of econometric procedures assessing the effect of an employment program on trainee earnings do not recover the experimental findings. This is considered to be the start of experimental benchmarking in social science.[235] | United States |
| 1986 | Concept | Fred N. Kerlinger describes the MAXMINCON principle, emphasizing maximizing systematic variance, controlling extraneous variance, and minimizing error variance in experimental design.[193] | United States |
| 1987 | Publication | Australian mathematician Anne Penfold Street and her daughter, statistician Deborah Street, publish Combinatorics of Experimental Design, a textbook connecting the combinatorial mathematics of block designs, Latin squares, and factorial designs to their applications in statistics. Reviewers praised the book for making the combinatorial side of experimental design accessible to statisticians, with Marshall Hall calling it "very readable" and "very satisfying," though some noted it omitted certain topics covered by more comprehensive contemporary texts.[236][237] | Australia |
| 1987 | Concept | S.K. Wang and Anastasios A. Tsiatis publish "Approximately optimal one-parameter boundaries for group sequential trials" in Biometrics, introducing a general parameterized family of stopping boundaries of which both the O'Brien–Fleming boundary (1979) and the Pocock boundary (1977) are shown to be special cases.[238] | United States |
| 1988 | Publication | Roger Mead publishes The Design of Experiments: Statistical Principles for Practical Applications, a textbook presenting practical principles of experimental design and analysis.[239] | United Kingdom |
| 1988 | Method | Organizational psychologists Miriam Erez and Gary P. Latham, in a dispute over the effect of participation on goal commitment and performance in goal setting research, design four experiments together with Edwin Locke serving as a neutral third party to resolve their disagreement — one of the earliest documented modern examples of what would later be termed adversarial collaboration, though the term itself was not yet coined.[240] | United States |
| 1988 | Organization | Stat-Ease releases its first version of Design–Expert, a statistical software package specifically dedicated to performing design of experiments (DOE), offering comparative tests, screening, characterization, and optimization tools.[241] | United States |
| 1989 | Experiment | The National Institute of Mental Health's Treatment of Depression Collaborative Research Program, led by Irene Elkin, publishes results comparing the tricyclic antidepressant imipramine, a placebo (each administered with supportive clinical management), and two forms of psychotherapy (cognitive-behavioral and interpersonal) for depression. The study finds no statistically significant differences in outcome across the four conditions overall, a result widely attributed to the unexpectedly strong performance of the placebo-plus-clinical-management condition — an influential demonstration of how effective a well-administered placebo control can be in psychotherapy and psychiatric drug research.[119][242] | United States |
| 1989 | Publication | Perry D. Haaland publishes Experimental Design in Biotechnology, presenting statistical experimental design and analysis as a problem-solving tool in biotechnology.[243][244][245] | United States |
| 1989 | Method | Jerome Sacks, William Welch, Toby Mitchell, and Henry Wynn publish a landmark paper summarizing the Bayesian statistical framework for computer experiments, modeling a deterministic computer simulation's output as an unknown function of its inputs and placing a Gaussian process prior over that function — establishing an approach to designing and analyzing simulation experiments distinct from classical experimental design for physical systems, where criteria like replication and A/D-optimality (suited to parametric models with random error) do not directly apply.[246] | United States |
| 1989 | Publication | Jerome Sacks and collaborators discuss statistical issues in the design and analysis of computer and simulation experiments, helping establish the field of computer experiments.[247] | United States |
| 1989 | Concept | C. W. H. Lam and collaborators complete a large-scale computer search establishing that no finite projective plane of order 10 exists — a case the Bruck–Ryser–Chowla theorem's necessary conditions do not themselves rule out, showing that theorem's criteria, while powerful, are not sufficient to determine existence in general. The proof, which combined coding theory with an extensive computer search, was notable enough to draw coverage in The New York Times questioning whether a computer-assisted proof that no human could fully verify by hand should count as a mathematical proof.[248][249] | Canada |
| 1990 | Organization | The International Council for Harmonisation (ICH) holds its inaugural meeting in Brussels, Belgium, hosted by the European Federation of Pharmaceutical Industries and Associations (EFPIA), bringing together regulatory authorities and pharmaceutical industry representatives from Europe, Japan, and the United States. The initiative aims to harmonize technical requirements for drug development and registration across jurisdictions, reducing redundant testing while ensuring the quality, safety, and efficacy of medicines. ICH would go on to produce more than 80 guidelines and expand to around 50 participating medicine authorities worldwide by its 30th anniversary in 2020.[250] | Global |
| 1993 | Concept | Judea Pearl introduces the Back-Door criterion, a graphical condition using causal graphical model for identifying a sufficient set of variables to adjust for in order to obtain an unbiased estimate of a causal effect — providing a formal, graph-based alternative to the counterfactual definitions of confounding developed in epidemiology, later shown to be formally equivalent to them.[251] | United States |
| 1993 | Policy | The Council for International Organizations of Medical Sciences (CIOMS) — founded in 1949 under the auspices of the World Health Organization and UNESCO — publishes International Ethical Guidelines for Biomedical Research Involving Human Subjects, a revision of its original 1982 guidelines prompted by the HIV/AIDS pandemic, proposed large-scale prevention and treatment trials, and the growing use of multinational field trials involving vulnerable populations. Though carrying no legal force, the guidelines become influential in shaping national approaches to research ethics review, and CIOMS would go on to revise them again in 2002 and 2016.[252] | Global |
| 1994 | Method | American psychiatrist Peter Breggin applies the challenge–dechallenge–rechallenge (CDR) protocol — administering, withdrawing, then re-administering a medication while monitoring for adverse effects — to investigate a suspected association between fluoxetine (Prozac) and suicidal ideation. Breggin had observed that only certain individuals responded to the medication with increased suicidal thoughts; given the low occurrence rate of this reaction, conventional statistical testing across a study population was considered inappropriate, making the individual-level CDR protocol a more suitable design. Eli Lilly and Company subsequently adopted the CDR protocol, rather than a randomized controlled trial, when testing for increased suicide risk associated with the drug.[253] | United States |
| 1994 | Experiment | David Card and Alan Krueger publish a study using the difference in differences design to evaluate the employment effects of New Jersey's April 1992 minimum wage increase, comparing fast-food employment in New Jersey against neighboring Pennsylvania (used as a control not subject to the wage increase) before and after the change. Contrary to standard economic theory's prediction that a minimum wage increase would reduce employment, they find no such decrease — an early, influential demonstration of using a natural experiment and a control region to estimate a causal treatment effect from observational data. Card received the 2021 Nobel Memorial Prize in Economic Sciences in part for this and related work.[254] | United States |
| 1994 | Method | The Neyer-d optimal test is first described by Barry T. Neyer.[255] | United States |
| 1995 | Method | Chris Nachtsheim and Ruth K. Meyer introduce the coordinate exchange algorithm in "The Coordinate-Exchange Algorithm for Constructing Exact Optimal Experimental Designs," published in Technometrics, enabling computational generation of optimal experimental designs. The cyclic coordinate-exchange algorithm constructs D-optimal and linear-optimal designs without needing to explicitly construct or enumerate candidate sets — which grow exponentially in the number of factors — and handles convex or mixed convex/discrete design spaces directly, without requiring sophisticated nonlinear programming routines. For design problems with 10 or more factors, the approach typically achieves reductions in computing time of two or more orders of magnitude compared to standard candidate-set-based procedures, with no loss of design efficiency.[256] | United States |
| 1995 | Concept | Leslie Kish proposes the "Design Effect Factor" (Deft), a refinement of his 1965 design effect that uses simple random sampling with replacement in the denominator rather than without, arguing this better isolates the effect of sampling design from the nuisance of finite-population correction and is simpler to use directly in confidence-interval calculations.[257] | United States |
| 1995 | Publication | Kathryn Chaloner and Isabella Verdinelli publish "Bayesian Experimental Design: A Review" in Statistical Science, a widely cited survey establishing the now-standard approach of assuming approximate normality of posterior probabilities in order to calculate expected utility using linear theory — becoming a standard reference point for later work in Bayesian experimental design.[258] | United States |
| 1996 | Ethical oversight | International Conference on Harmonisation (ICH) guidelines for Good Clinical Practice (GCP) established.[259] | Global |
| 1996 | Method | Colombian-Canadian physician Alex Jadad, working as a Research Fellow at Oxford's Pain Relief Unit, and colleagues introduce the Jadad scale (also known as Jadad scoring or the Oxford quality scoring system) in an appendix to a paper on blinding in randomized clinical trials. The instrument scores a trial report from zero to five based on three yes/no questions covering randomization, double-blinding, and reporting of withdrawals and dropouts. Despite later criticism for over-emphasizing blinding, showing low inter-rater consistency, and omitting allocation concealment, the scale becomes the most widely used trial-quality assessment tool worldwide, with its seminal paper cited in over 25,000 scientific works by 2024.[260] | United Kingdom |
| 1996 | Method | Stat-Ease releases Design–Expert Version 5, the first version of the software designed for Microsoft Windows, broadening accessibility of dedicated design-of-experiments software beyond earlier DOS-based tools.[261] | United States |
| 1996 | Policy | The Consolidated Standards of Reporting Trials (CONSORT) Statement is first published, the product of a 1995 Chicago meeting merging two independent efforts to improve randomized-trial reporting: the 1993 Ottawa meeting's Standardized Reporting of Trials (SORT) proposal and the concurrent Asilomar Working Group's recommendations from California. Convened at the suggestion of JAMA's Drummond Rennie, the merged CONSORT Statement provides a standardized checklist and participant flow diagram intended to reduce bias and aid critical appraisal of trial reports.[262] | United States |
| 1997 | Concept | German researchers Gunver Kienle and Helmut Kiene publish a reanalysis of Henry K. Beecher's influential 1955 paper "The Powerful Placebo" in the Journal of Clinical Epidemiology, re-examining the original studies Beecher had cited as evidence of a roughly 35% placebo response rate. They report finding no genuine evidence of a placebo effect in any of the studies Beecher relied upon, attributing his conclusions instead to factors such as natural disease progression, statistical regression to the mean, and co-administered treatments — a finding that substantially undercuts the empirical foundation of the "powerful placebo" narrative that had gone largely unchallenged for four decades.[263] | Germany |
| 1997 | Concept | Computer scientist Tom M. Mitchell proposes a widely cited formal definition of machine learning: "A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P if its performance at tasks in T, as measured by P, improves with experience E." — offering an operational, experimentally-grounded definition rather than an aspirational one, directly framing learning in terms of measurable performance under controlled tasks.[264] | United States |
| 1998 | Policy | The U.S. Department of Health and Human Services announces a policy change to the population standard used for age-adjusting death rates in its publications.[265] | United States |
| 1998 | Experiment | The UK Prospective Diabetes Study Group publishes UKPDS 38 in the BMJ, reporting results from a trial comparing tight blood pressure control (<150/85 mmHg) against less strict control (<180/105 mmHg) in patients with type 2 diabetes. The trial is stopped before its planned completion because tight control proves so much more effective at preventing macrovascular and microvascular complications — including death, myocardial infarction, and stroke — that continuing to withhold it from the control group is no longer considered ethical.[266] | United Kingdom |
| 1999 | Organization | The first International Data Farming Workshop is held at the Maui High Performance Computing Center, the start of a series that would go on to meet roughly twice a year for decades, eventually organized by the SEED Center for Data Farming at the Naval Postgraduate School. The first four workshops focus on methodology — complex adaptive systems modeling, agent-based representation, and statistical experiment design — before subsequent workshops shift toward applications, with participation from countries including Canada, Singapore, Mexico, Turkey, and the United States assigning teams to apply data farming techniques to domains such as robotics, homeland security, and disaster relief. By the 30th workshop in February 2016, data farming — a term Gary Horne had coined in 1997 — had become an established methodology within NATO's modeling and simulation community.[267] | United States |
| 1999 | Method | Rajeev Dehejia and Sadek Wahba re-examine Robert LaLonde's original 1986 dataset using additional non-experimental methods, arguing that when there is sufficient overlap between treated and untreated subject pools and unobservable covariates do not substantially impact outcomes, non-experimental methods can in fact estimate treatment effects accurately — offering a more optimistic counterpoint to LaLonde's original benchmarking critique.[268] | United States |
| 1999 | Policy | The oral rotavirus vaccine RotaShield, licensed in the United States in 1998, is withdrawn less than a year later after post-marketing Phase IV surveillance identifies an elevated risk of intussusception (a form of bowel obstruction) among vaccinated infants. The episode becomes a widely cited case study illustrating the necessity of Phase IV surveillance for detecting rare adverse events that trials of the size typically used in Phase III cannot reliably capture.[269] | United States |
| 1999 | Concept | Basili et al use the term family of experiments to refer to a group of experiments that pursue the same goal and whose results can be combined into joint—and potentially more mature—findings than those that can be achieved in isolated experiments.[270] | United States |
| 2000 (January 19) | Publication | A First Course in Design and Analysis of Experiments.[271] | United States |
| 2001 | Method | The World Health Organization publishes a new standard population for age standardization, intended to allow health statistics from different countries to be compared despite differing population age profiles.[272] | Global |
| Late 20th century | Concept | Fisher’s principles of randomization, replication, and blocking become standard features of statistically rigorous experiments in the biological and biomedical sciences.[18] | Global |
| Late 20th century | Concept | Recognition that all experiments are inherently designed, with emphasis on the importance of proper planning to avoid wasted resources and invalid results.[20] | Global |
| Late 20th century | Method | Advances in computing and algorithms enable practical implementation of optimal design, expanding its use in scientific and industrial applications.[20] | Global |
| 21st century | Application | Designed experiments are widely applied across service sectors including finance, business operations, and government, reflecting the broad adoption of statistical experimentation.[21] | Global |
| 2001 | Method | Daniel Kahneman independently develops a protocol for adversarial collaboration roughly a decade after Erez, Latham, and Locke's earlier unnamed example, and may have been the first to use the term itself.[273][274] | United States |
| 2001 | Concept | A study led by P.J. Devereaux, published in JAMA, surveys practicing physicians and reviews textbook definitions of "single-blind," "double-blind," and "triple-blind" terminology in randomized controlled trials, finding substantial disagreement both among physicians and between physicians and published textbook definitions as to which parties (patients, treating clinicians, data collectors, outcome assessors, data analysts) each term is understood to cover. The study becomes an early, influential demonstration — predating Haahr and Hróbjartsson's 2006 trial-level analysis of the same problem — that "blinding" terminology is used inconsistently enough to obscure rather than clarify a trial's actual methodology.[275] | Canada |
| 2001 | Policy | Following further CONSORT Group meetings in 1999 and 2000, a revised CONSORT Statement is published, updating the original 1996 recommendations for reporting parallel-group randomized trials in light of growing empirical evidence on reporting quality.[276] | Global |
| 2002 | Policy | The World Medical Association issues a "Note of Clarification" on Paragraph 29 of the Declaration of Helsinki, addressing controversy over the ethics of placebo-controlled trials. The note reaffirms that placebo-controlled methodology should generally be used only in the absence of an existing proven therapy, but carves out two exceptions where it may remain ethically acceptable even when proven therapy exists: when compelling and scientifically sound methodological reasons make a placebo necessary to determine efficacy or safety, or when the condition under investigation is minor and placebo recipients face no additional risk of serious or irreversible harm. The clarification becomes a widely cited reference point in ongoing debates over the ethics of withholding proven treatment from trial participants.[277] | Global |
| 2002 | Method | Jochen Musch and Karl Christoph Klauer, and separately Ulf-Dietrich Reips, describe the "seriousness check" for web-based psychological experiments: asking respondents at the start of an online study whether they intend to participate seriously or merely browse the pages, in order to identify and exclude low-quality data entries before analysis. Later research finds the technique to be a strong predictor of dropout — roughly 75% of respondents who say they only want to look at the pages subsequently drop out, versus only 10–15% of those who say they intend to seriously participate — and that overall 30–50% of visitors to a typical online study fail the check.[278][279] | Germany |
| 2002 | Method | Howard Bloom, Charles Michalopoulos, Carolyn Hill, and Ying Lei conduct a large-scale experimental benchmarking study of mandatory welfare-to-work programs, testing which non-experimental methods come closest to recovering experimentally estimated program effects. They conclude that none of the non-experimental methods tested approach the accuracy of a randomized experiment for recovering the parameter of interest, reinforcing concerns first raised by Robert LaLonde's 1986 benchmarking study.[280] | United States |
| 2002 | Concept | The terms exploratory thought and confirmatory thought are introduced by social psychologist Jennifer Lerner and psychology professor Philip Tetlock in their book Emerging Perspectives in Judgment and Decision Making.[281] | United States |
| 2003 | Organization | The Abdul Latif Jameel Poverty Action Lab is founded to scale randomized evaluations in development economics.[282] | United States |
| 2003 | Concept | Economist Colin Camerer coins the term "experimetrics" in his book Behavioral Game Theory: Experiments in Strategic Interaction, to describe econometric techniques customized for the analysis of experimental economic data — ranging from simple treatment comparisons to complex multi-equation structural models. The term comes to describe the broader intersection of experimental economics and econometrics, encompassing both the statistical methodology for analyzing experimental data and, less commonly, the optimal design of the experiments themselves.[283] | United States |
| 2003 | Method | Japanese researcher S. Hirata designs the loose-string task, which becomes the standard apparatus for cooperative pulling experiments: a single string threaded through loops on a movable platform such that if only one participant pulls, the string comes loose and the reward becomes unretrievable, requiring genuinely coordinated pulling for success. This design, simpler and more broadly applicable across species than Crawford's original box-and-rope apparatus, is subsequently adopted in cooperative pulling studies of rooks, ravens, wolves, elephants, capuchins, and many other species.[284] | Japan |
| 2003 | Method | Alberto Abadie and Javier Gardeazabal introduce the synthetic control method in "The Economic Costs of Conflict: A Case Study of the Basque Country," published in the American Economic Review. The method estimates a counterfactual for a single treated unit (such as a region or country) by constructing a "synthetic" control as a data-driven weighted average of untreated comparison units, chosen so the weighted combination closely matches the treated unit's pre-intervention characteristics and outcome trajectory. Unlike difference in differences, the approach can account for confounders whose effects change over time.[285] | Spain |
| 2004 | Policy | The U.S. Food and Drug Administration publishes the report "Innovation or Stagnation: Challenge and Opportunity on the Critical Path to New Medical Products," calling attention to a "slowdown... in innovative medical therapies reaching patients" and identifying areas of drug development needing improvement, including predictive models, biomarkers, and new clinical evaluation techniques. Through the resulting Critical Path Initiative, FDA works to foster innovation via collaborations among government, industry, academia, and patient advocacy groups; adaptive designs for clinical trials initially emerge under this regulatory framework.[286] | United States |
| 2004 | Policy | The CONSORT group publishes an extension to the CONSORT statement specifically for cluster randomised trials, addressing the additional reporting requirements these designs need — such as accounting for intraclass correlation — beyond standard individually randomised trial reporting guidelines.[287] | United Kingdom |
| 2005 | Experiment | Study determines that most clinical trials have unclear allocation concealment in their protocols, in their publications, or both.[288] | United Kingdom |
| 2005 | Publication | Stuart Pocock publishes the editorial "When (not) to stop a clinical trial for benefit" in JAMA, addressing the range of practical and ethical considerations bearing on the decision to halt a trial early when interim results favor the treatment group — beyond the purely statistical threshold his own Pocock boundary had provided in 1977 — including the risk that early, promising results may not persist or generalize, and the tension between statistical stopping rules and the ethical pressure to give patients access to an apparently superior treatment.[289] | United Kingdom |
| 2005 | Concept | John Ioannidis publishes "Why Most Published Research Findings Are False" in PLOS Medicine, arguing on theoretical and probabilistic grounds that a majority of published research claims across many fields are likely to be false, due to factors including small sample sizes, small effect sizes, flexibility in study design and analysis, and bias from financial or other interests. The paper becomes one of the most-cited and most-discussed articles in the history of the reproducibility crisis debate, predating and helping motivate the preregistration and open-science reforms of the following decade.[290] | United States |
| 2005 | Publication | Jack Kleijnen, Susan Sanchez, Thomas Lucas, and Thomas Cioppa publish "A User's Guide to the Brave New World of Designing Simulation Experiments" in INFORMS Journal on Computing, addressing a key limitation of early data farming work: initial reliance on brute-force full factorial designs meant only a small number of factors could be investigated due to the curse of dimensionality. The paper helps establish improved experimental designs specifically suited to large-scale simulation experiments, developed in collaboration between Project Albert and the Naval Postgraduate School's SEED Center for Data Farming.[291] | United States |
| 2006 | Publication | American psychologist Seth Roberts publishes The Shangri-La Diet, a popular diet book based on conclusions he drew from self-experimentation — informally applying the N-of-1 trial logic Allen Neuringer had proposed for self-experimentation in 1981. Roberts, who documented his self-experiments on his blog, becomes a prominent early figure associated with the later quantified self movement, in which the growing ease of personal data collection and analysis drives a proliferation of N-of-1-style personal experiments.[292] | United States |
| 2006 | Policy | The U.S. Food and Drug Administration issues its Guidance on Exploratory Investigational New Drug (IND) Studies, formally introducing "Phase 0" trials — optional, exploratory human microdosing studies that administer a single subtherapeutic dose of a candidate drug or imaging agent to a small number of subjects (10–15) to gather early pharmacokinetic data. Because doses are too low to produce any therapeutic effect, Phase 0 trials yield no safety or efficacy data, but let developers rank candidate drugs and make go/no-go decisions based on relevant human data rather than solely on animal models, before committing to a full Phase I trial. The designation, initially unusual, becomes generally adopted as standard practice in drug development.[293] | United States |
| 2006 | Concept | Danish researchers Mette T. Haahr and Asbjørn Hróbjartsson examine a random sample of 200 randomized controlled trials described as "double blind" and survey their authors, finding that blinding of patients, care providers, and assessors was clearly and separately described in only three trials (2%), while 56% failed to describe the blinding status of any individual involved at all. The study demonstrates that labelling a trial "double blind" does not reliably indicate which parties were actually blinded, undermining the term's usefulness as reported in trial literature.[294] | Denmark |
| 2006 | Method | Alicia Melis, Brian Hare, and Michael Tomasello design a cooperative pulling experiment that explicitly controls for social tolerance between partners — a variable earlier studies had not accounted for — by comparing captive chimpanzee pairs known to share food readily against pairs less inclined to do so. They find food-sharing tolerance strongly predicts cooperative success, resolving much of the inconsistency across earlier chimpanzee cooperative pulling studies and establishing social tolerance as a key confounding factor that subsequent cooperative pulling experiments across many species would need to control for.[295] | Germany |
| 2007 (1 April) | Organization | The National Research Ethics Service (NRES) launches in the United Kingdom, a body requiring principal investigators to obtain approval for proposed research studies involving human participants before proceeding — with unapproved studies prohibited. NRES describes its purpose as reviewing research proposals "to protect the rights and safety of research participants and enable ethical research which is of potential benefit to science and society," extending the institutional ethics-committee review model established internationally by the Declaration of Helsinki decades earlier into a dedicated national service. The word "National" is later dropped from the name at an unrecorded point, and its functions are absorbed into the NHS Health Research Authority's Research Ethics Service.[296] | United Kingdom |
| 2008 | Method | Economist Justin McCrary proposes the "density test" for regression discontinuity designs, examining whether the density of observations of the assignment variable is continuous at the treatment cutoff. A discontinuity in this density — for example, an unusually large number of students who "just barely" pass an exam relative to those who "just barely" fail — suggests that some participants may have been able to manipulate their treatment status, undermining the "as good as random" assumption the design's validity depends on. The test becomes a standard diagnostic check in applied regression discontinuity research.[297] | United States |
| 2008 | Method | Bradley Jones, Dennis K. J. Lin, and Chris Nachtsheim introduce Bayesian D-optimal supersaturated designs in the Journal of Statistical Planning and Inference, a new class of supersaturated designs that can accommodate arbitrary sample sizes, blocks of any size, and categorical factors with more than two levels — matching or outperforming the best previously published designs on the E(s²) criterion for two-level experiments with even sample size.[298] | United States |
| 2008 | Concept | A meta-epidemiological study of 146 meta-analyses finds that randomized controlled trials with inadequate or unclear allocation concealment tend to show results biased toward beneficial treatment effects — but only when trial outcomes are subjective rather than objective, refining earlier, less qualified claims about the relationship between allocation concealment and bias.[299] | United Kingdom |
| 2009 | Method | Adversarial collaboration is recommended by Daniel Kahneman[300] and others as a way of resolving contentious issues in fringe science, such as the existence or nonexistence of extrasensory perception.[301] | United States |
| 2010 | Concept | Asbjørn Hróbjartsson and Peter C. Gøtzsche argue in a meta-analysis that observed placebo effects can arise from bias due to lack of blinding, emphasizing the importance of proper control and blinding in experimental design.[302] | Denmark |
| 2010 | Policy | The National Centre for the Replacement, Refinement and Reduction of Animals in Research (NC3Rs) publishes the ARRIVE guidelines (Animals in Research: Reporting In Vivo Experiments), a 20-item checklist for improving experimental design and reporting standards in animal research, modeled on the CONSORT statement for reporting randomized trials. The guidelines follow a 2009 NC3Rs review finding that most biomedical journals provided little guidance on animal research design and reporting, and that large proportions of published animal studies failed to state a hypothesis, describe animal characteristics, use randomization or blinding, or fully report statistical methodology.[303][304] | United Kingdom |
| 2010 | Policy | The U.S. Food and Drug Administration issues draft guidance on adaptive design for clinical trials, an early regulatory step preceding the agency's more comprehensive 2019 guidance.[305] | United States |
| 2010 | Method | Researchers at the European Bioinformatics Institute (EMBL-EBI), led by J. Malone, publish a paper describing the Experimental Factor Ontology (EFO), an open-access ontology for systematically modeling experimental and sample variables — covering disease, anatomy, cell type, cell lines, chemical compounds, and assay information — originally developed to describe experimental variables in EBI's Expression Atlas resource. EFO is built to interoperate with existing biomedical ontologies such as ChEBI and the Ontology for Biomedical Investigations, and becomes a cross-cutting resource used for curation, querying, and data integration across resources including Ensembl and ChEMBL.[306] | United Kingdom |
| 2010 | Method | Alberto Abadie, Alexis Diamond, and Jens Hainmueller formalize and popularize the synthetic control method in "Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California's Tobacco Control Program," published in the Journal of the American Statistical Association, applying the technique to estimate the causal effect of California's 1988 tobacco-control program by constructing a synthetic "California" from a weighted combination of other U.S. states. The paper establishes synthetic control as a systematic, data-driven alternative to ad hoc comparison-group selection in policy evaluation, and the method sees subsequent application across economics, political science, health policy, criminology, and drug development.[307] | United States |
| 2010 | Policy | The CONSORT 2010 Statement is published, consisting of a 25-item checklist and participant flow diagram, alongside an "Explanation and Elaboration" document detailing the reasoning behind each recommendation. By this point the CONSORT Statement has been endorsed by over 600 journals and editorial groups, including The Lancet, BMJ, JAMA, and the New England Journal of Medicine, and has directly inspired parallel reporting-guideline initiatives for other study types, including STROBE for observational studies and PRISMA for systematic reviews.[308] | Global |
| 2010 | Experiment | The I-SPY 2 trial launches as an adaptive Phase 2 clinical trial platform for breast cancer, linking academic cancer centers, the FDA, the NCI, and pharmaceutical partners. Building on predictive biomarkers developed in its predecessor I-SPY 1 (2002–2006), the trial evaluates multiple experimental drug combinations against standard chemotherapy simultaneously, dropping ineffective regimens early and advancing promising ones to confirmatory trials — a widely cited real-world demonstration of adaptive experimental design accelerating drug development.[309] | United States |
| 2011 | Method | Bradley Jones, Principal Research Fellow for the JMP Division of SAS, and Chris Nachtsheim, Chair of the Operations and Management Science Department at the University of Minnesota's Carlson School of Management, publish "A Class of Three-Level Designs for Definitive Screening in the Presence of Second-Order Effects" in the Journal of Quality Technology, developing definitive screening designs — three-level designs providing main-effect estimates unbiased by any second-order effect, requiring only one more than twice as many runs as factors, while avoiding confounding between any pair of second-order effects.[310] | United States |
| 2011 | Method | Stefano M. Iacus, Gary King, and Giuseppe Porro introduce Coarsened Exact Matching (CEM) in the Journal of the American Statistical Association, a matching method they present as simpler, more statistically powerful, and monotonic imbalance bounding compared to earlier techniques such as propensity score matching.[311] | United States |
| 2013 | Method | Alex Deng, Ya Xu, Ron Kohavi, and Toby Walker introduce CUPED (Controlled-experiment Using Pre-Experiment Data) at the WSDM international conference, a technique for improving the statistical sensitivity of online A/B tests by using data from before an experiment began to reduce the variance of outcome measurements — allowing web companies to detect smaller effects with the same sample size, or the same effects with less traffic. The technique becomes widely adopted in large-scale online experimentation at companies running many concurrent experiments on millions of users.[312] | United States |
| 2014 | Concept | Uri Simonsohn, Leif Nelson, and Joseph Simmons — the researchers behind the blog Data Colada — coin the term "p-hacking" to describe the practice of running many statistical analyses on the same data set and reporting only those that yield a statistically significant result, dramatically understating the true risk of false positives. The paper also introduces the "p-curve" technique for detecting p-hacking in published literature by examining the distribution of significant p-values across studies.[313] | United States |
| 2014 | Concept | Peter Keevash proves the existence of nontrivial Steiner systems for all t ≥ 6, resolving a long-standing open problem in design theory, in "The existence of designs." His proof is non-constructive, and as of 2019 no explicit Steiner systems are actually known for large values of t despite the existence proof.[314] | United Kingdom |
| 2014 | Concept | Andrew Gelman and John Carlin (statistician) propose Type S (sign) and Type M (magnitude) errors as complements to the traditional Type I/Type II framework, in "Beyond power calculations: Assessing Type S (sign) and Type M (magnitude) errors," published in Perspectives on Psychological Science. A Type S error occurs when a statistically significant result's estimated effect has the wrong sign relative to the true effect — a risk that grows in low-powered studies — while a Type M error concerns how much a significant result's estimated effect size is inflated relative to the true effect, a consequence of selection bias introduced by using significance itself to filter which results get reported.[315] | United States |
| 2014 | Method | Researchers at McMaster University led by Michael Walsh introduce the Fragility Index (FI) in "The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index," published in the Journal of Clinical Epidemiology. The FI measures the minimum number of participants in a trial whose outcome would need to switch from a non-event to an event (or vice versa) to change a statistically significant result to a non-significant one, exposing how easily the conclusions of many published, adequately-powered randomized controlled trials can be reversed by the reclassification of only one or a few patients — a concern the authors argue is especially acute for trials with small sample sizes and few observed events.[316] | Canada |
| 2015 | Experiment | Journalist John Bohannon deliberately conducts and publishes a fraudulent study claiming chocolate consumption accelerates weight loss, using p-hacking techniques — considering 18 different variables during testing until one produced a statistically significant result — to demonstrate publicly how easily scientific-sounding claims can be manufactured from real but hollow statistical analysis. The hoax was picked up uncritically by numerous media outlets before Bohannon revealed it was an intentional social experiment exposing weaknesses in science journalism and nutrition research.[317] | Germany |
| 2015 | Concept | Megan L. Head, Luke Holman, Rob Lanfear, Andrew T. Kahn, and Michael D. Jennions publish "The Extent and Consequences of P-Hacking in Science" in PLOS Biology, analyzing the distribution of p-values reported across a large sample of published papers to empirically estimate how widespread p-hacking actually is in scientific practice, rather than treating it as a purely theoretical risk. The study also examines factors — including pressure for early stopping mandated by some animal ethics boards when interim results reach significance — that can inadvertently encourage the practice.[318] | Australia |
| 2016 | Concept | A study finds that symphony orchestra auditions conducted behind a curtain, blinding judges to a performer's gender, increase the hiring of women — a widely cited real-world example of blinding used outside clinical or laboratory settings.[319] | United States |
| 2016 | Method | Susan Athey and Guido Imbens publish "Recursive partitioning for heterogeneous causal effects" in PNAS, pioneering machine learning techniques — adapting decision-tree methods to causal inference — for detecting and characterizing how treatment effects vary across subpopulations, addressing the external validity question of whether a treatment's effect generalizes across different subsets of people, times, and contexts rather than assuming a single homogeneous effect. Stefan Wager and Athey extend the approach using random forests in a 2018 follow-up paper.[320] | United States |
| 2017 | Policy | Nature Human Behaviour launches with its inaugural issue offering authors the option of publishing a registered report — a format in which the study proposal is peer-reviewed and receives in-principle acceptance before the research outcomes are known, guaranteeing publication regardless of result. The journal's founding editorial frames the format as a way to neutralize publication bias and remove incentives for practices that undermine the validity of research, shifting editorial evaluation from a study's results to the soundness of its questions and methods.[321] | United Kingdom |
| 2017 | Organization | The International Collaborative Network for N-of-1 Trials and Single-Case Designs (ICN) is established, co-chaired by Jane Nikles and Suzanne McDonald, as a global network of clinicians, researchers, and consumers interested in N-of-1 trials and single-case experimental designs; it grows to over 400 members across more than 30 countries.[322] | Global |
| 2017 | Method | Peter M. Aronow and Cyrus Samii publish "Estimating average causal effects under general interference, with application to a social network experiment" in the Annals of Applied Statistics, developing a general method for computing exposure probability matrices in experiments where SUTVA (the stable unit treatment value assumption) does not hold — such as experiments conducted over social networks where treated units can influence untreated neighbors — enabling inverse probability weighting approaches, including the Horvitz–Thompson estimator, to correct for the resulting bias when estimating average treatment and spillover effects.[323] | United States |
| 2017 | Concept | A historical case study documents the subversion of allocation concealment in a randomized controlled trial, part of a broader body of evidence that sealed-envelope and even centralized allocation concealment methods remain vulnerable to circumvention by study personnel — including researchers opening envelopes prematurely, holding them up to light, or keeping lists of previous allocations (reported by up to 15% of study personnel surveyed).[324] | United Kingdom |
| 2018 | Concept | A meta-epidemiological study of 540 randomized controlled trials in oral health research, led by Humam Saltaji, formally critiques the terms "double-blind" and "triple-blind," noting that blinding can be implemented at multiple distinct levels of a trial — participants, outcome assessors, care providers, data analysts, and investigators — and that "double blinding" or "triple blinding" may refer to any two or three of these levels without specifying which, producing "conceptual and operational ambiguity." The study finds that two-thirds of the trials examined were not in fact described as double-blind at all, and recommends, in line with CONSORT guidance, that researchers abandon the "double"/"triple" blind terminology altogether in favor of explicitly stating who was blinded and to which components of the trial.[325] | Canada |
| 2018 | Experiment | Brett Gordon, Florian Zettelmeyer, Neha Bhargava, and Dan Chapsky use data from large-scale field experiments on Facebook advertising to test whether standard observational methods — including propensity score matching, stratification, and regression adjustment — can recover the true causal effects of online ads on checkout, registration, and page-view outcomes. Despite the unusually rich variation available in social-media advertising data, they find observational methods are unable to accurately recover the causal effects established by the underlying randomized experiments, providing a striking contemporary demonstration of the same experimental-benchmarking concerns first raised by LaLonde in 1986.[326] | United States |
| 2018 | Concept | A study by Brian Nosek and colleagues proposes preregistration as a safeguard against p-hacking: researchers submit their data-analysis plan to a journal before beginning data collection, precluding after-the-fact manipulation of the analysis to reach statistical significance.[327] | United States |
| 2018 | Experiment | In a satirical study published in the BMJ's annual Christmas issue, Robert Yeh and colleagues randomize 23 volunteers to jump from an aircraft wearing either a parachute or an empty backpack, finding no significant difference in death or major injury between the two groups — a result made possible only because the aircraft was parked on the ground and participants jumped roughly two feet. The trial becomes a widely cited illustration, aimed at critics of evidence-based medicine's reliance on randomized controlled trials over clinical judgment, of how a methodologically valid RCT can nonetheless produce a misleading or non-generalizable conclusion when applied to a scenario where the intervention's benefit is already self-evident, and of the necessity of peer review in contextualizing trial results.[328] | United States |
| 2018 | Concept | Kenneth Schulz and colleagues publish a retrospective account of the introduction and adoption of the term "allocation concealment," distinguishing it from the related but distinct concept of blinded experiment: allocation concealment prevents foreknowledge of treatment assignment before randomization (addressing selection bias), while blinding conceals group identity after allocation (addressing ascertainment bias). The authors argue the earlier common term "randomization blinding" had confusingly conflated these two distinct sources of bias.[329] | United States |
| 2019 | Recognition | Nobel Prize in Economics awarded to Abhijit Banerjee, Esther Duflo, and Michael Kremer for experimental approaches to alleviating global poverty.[330] | Sweden |
| 2019 | Policy | The U.S. Food and Drug Administration provides guidance on the use of adaptive designs in clinical trials, formalizing regulatory standards for flexible experimental designs.[331] | United States |
| 2019 | Concept | Gary King and Richard Nielsen publish "Why Propensity Scores Should Not Be Used for Matching" in Political Analysis, arguing that propensity score matching — developed by Rubin and Rosenbaum in 1983 and widely adopted since — actually increases model dependence, bias, and inefficiency relative to other matching methods, and should no longer be the default recommendation. The paper marks a significant reversal in applied best practice among political scientists, economists, and other users of observational causal-inference methods.[332] | United States |
| 2019 | Concept | Aleksander Fabijan, Jayant Gupchup, Somit Gupta, Jeff Omhover, Wen Qin, Lukas Vermeer, and Pavel Dmitriev publish "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments" at the ACM SIGKDD conference, formalizing sample ratio mismatch (SRM) — a statistically significant discrepancy between the expected and actual ratio of treatment and control group sizes in an experiment, often caused by failures in randomization or measurement instrumentation in online A/B testing — and proposing a taxonomy and detection rules of thumb, typically using a chi-squared goodness-of-fit test, to help practitioners identify SRM and avoid drawing conclusions from biased experimental data.[333] | United States |
| 2019 | Method | Bradley Jones, Ryan Lekivetz, Dibyen Majumdar, Christopher J. Nachtsheim, and Jonathan W. Stallrich publish "Construction, Properties, and Analysis of Group-Orthogonal Supersaturated Designs" in Technometrics, developing group-orthogonal supersaturated designs (GO-SSDs) that improve factor screening when the number of variables exceeds the number of experimental runs.[334] | United States |
| 2020 | Experiment | In response to the COVID-19 pandemic, the World Health Organization launches the Solidarity trial, European researchers launch the Discovery trial, and UK researchers launch the RECOVERY Trial — each a large-scale, multi-arm adaptive design clinical trial of candidate treatments for hospitalized patients with severe COVID-19 infection. The adaptive designs allow ineffective experimental treatments to be dropped quickly and replaced with others as evidence accumulates, letting researchers adjust trial parameters in near real time rather than waiting for a trial's predetermined endpoint, and become widely cited examples of adaptive trial methodology deployed at unprecedented speed and scale during a public health emergency.[335][336] | Global |
| 2020 | Policy | An international working group supported by the National Centre for the Replacement, Refinement and Reduction of Animals in Research publishes ARRIVE 2.0, a revision of the 2010 ARRIVE guidelines splitting the original 20-item checklist (effectively 38 items counting sub-items) into an "Essential 10" checklist of basic minimum requirements and a "Recommended Set" of 11 additional items, in response to studies finding the original guidelines — despite endorsement by over 600 journals by 2016 and over 1,000 by 2020 — had been largely ignored by researchers and made little measurable impact on reporting quality.[337][338] | United Kingdom |
| 2020 | Method | Researchers Elias Garcia-Pelegrin, Alexandra Schnell, Clive Wilkins, and Nicola Clayton argue in Science that magic tricks offer a novel approach to hypothesis testing and experimental design in the study of animal cognition: presenting a trick that reliably fools humans to a nonhuman animal, then using its behavioral response (e.g. prolonged looking time, taken as an indicator of surprise) to probe which cognitive blind spots — such as expectations about object permanence or attention control — the animal shares with humans. The approach faces practical challenges distinct from conventional experimental designs, including getting animals to attend to a human demonstrator and inferring "surprise" without verbal report.[339] | United Kingdom |
| 2020 | Concept | Jérôme Adda, Christian Decker, and Marco Ottaviani publish "P-hacking in clinical trials and how incentives shape the distribution of results across phases" in PNAS, providing empirical evidence that the distribution of statistically significant results in clinical trials shifts in ways consistent with p-hacking as financial and career incentives change across different phases of drug development — connecting the general phenomenon of p-hacking directly to the economics of pharmaceutical research.[340] | Italy |
| 2020 | Concept | Gonville and Caius College, Cambridge removes a stained-glass window honoring Ronald Fisher — which had depicted a 7×7 Latin square in reference to his Design of Experiments — because of Fisher's association with eugenics. The removal, following a vote by the college's governing body, becomes a widely covered episode in the broader reckoning of statistics and genetics with the eugenic commitments of several of the field's founding figures.[341] | United Kingdom |
| 2021 | Method | Researchers Marianne Gunderson, Kristian A. Bjørkelo, and Jill Walker Rettberg, together with a team of larp designers led by Anita Myhre Andersen, stage Sivilisasjonens venterom ("Civilization's Waiting Room") in Bergen, Norway — a live-action roleplaying game (larp) designed as a research methodology in its own right, intended to let participants practice ethical decision-making around emerging surveillance and machine-vision technologies (facial recognition, deepfakes, VR) within a fictional AI-governed society. The project is analyzed in subsequent scholarship as a "mimetic method" related to design fiction, illustrating how immersive, participatory fiction can function as a qualitative research design distinct from conventional controlled experimentation.[342] | Norway |
| 2021 | Experiment | Garcia-Pelegrin, Schnell, Wilkins, and Clayton perform three sleight-of-hand magic tricks — palming, the French drop, and fast pass — on six Eurasian jays, an early empirical test of the magic-based experimental paradigm they had proposed the previous year. The jays are not deceived by palming or the French drop, both of which depend on human hand movements setting expectations about object location, but are successfully deceived by the faster "fast pass" technique — the first demonstrated case of a magic trick fooling a nonhuman animal.[343] | United Kingdom |
| 2022 | Organization | Philip Tetlock and Cory Clark propose adversarial collaboration as a vehicle for scientific self-correction, arguing it helps expose false claims by forcing exploration of rival hypotheses; their work leads to the University of Pennsylvania School of Arts & Sciences establishing the Adversarial Collaboration Project to formally support and encourage the approach across research questions.[344][345] | United States |
| 2024 | Concept | Cochrane (organisation) publishes an updated meta-epidemiological review by Ingrid Toews and colleagues comparing healthcare outcomes assessed in observational studies against those assessed in randomized controlled trials, finding little evidence of significant effect differences between the two designs regardless of the specific observational design used — a notably more optimistic finding for observational methods than the experimental-benchmarking critiques raised by Robert LaLonde (1986) and others in economics and social science.[346] | Global |
| 2025 | Policy | The CONSORT Group issues the CONSORT 2025 checklist, superseding the 2010 statement by adding seven new items, revising three, removing one, and introducing a new open-science reporting section covering trial registration, protocol and statistical-analysis-plan access, data sharing, and conflict-of-interest disclosures.[347] | Global |
| 2025 | Concept | Physician-scientist Pavlos Msaouel publishes "The curious rise of randomised non-comparative trials" in Significance, critiquing a design — traced by a companion meta-epidemiological review to oncology studies dating back to 2002 — in which participants are randomized to different treatment arms, but each arm is compared only against a historical control or predefined benchmark rather than against each other, functioning in practice as multiple concurrent single-arm studies. Msaouel argues the term "randomisation" in these designs is largely "talismanic," misleadingly suggesting the methodological benefits of a true randomized controlled trial despite providing no formal between-arm comparison; a companion analysis by Alexander D. Sherry, Msaouel, and Ethan B. Ludmir finds that despite this, roughly half of published randomised non-comparative trials still report some form of comparison between arms anyway.[348][349] | United States |
| 2025 | Policy | Miguel Hernán, Aidan G. Cashin, and an international group of collaborators publish the TARGET Statement in JAMA, a reporting guideline for observational studies that use the "target trial emulation" framework — designing an observational study to explicitly emulate the randomized controlled trial that would ideally answer the same causal question, in order to reduce common biases in observational causal inference. The statement extends the lineage of trial-reporting guidelines established by CONSORT (for randomized trials) and STROBE (for observational studies generally) to this specific and increasingly widely used design.[350] | United States |
| 2020s | Method | Machine learning is integrated with experimentation to estimate heterogeneous treatment effects and optimize interventions | Global [351] |
Numerical and visual data
Google Scholar
The following table summarizes per-year mentions on Google Scholar as of December 14, 2021.
| Year | "experimental design" |
|---|---|
| 1900 | 30 |
| 1910 | 17 |
| 1920 | 13 |
| 1930 | 19 |
| 1940 | 62 |
| 1950 | 425 |
| 1960 | 1,590 |
| 1970 | 6,240 |
| 1980 | 11,400 |
| 1990 | 17,000 |
| 2000 | 53,200 |
| 2010 | 162,000 |
| 2020 | 90,600 |

Google Trends
The chart below shows Google Trends data for Design of experiments (Topic), from January 2004 to December 2021, when the screenshot was taken. Interest is also ranked by country and displayed on world map.[352]

Google Ngram Viewer
The chart below shows Google Ngram Viewer data for Design of experiments, from 1900 to 2019.[353]

Wikipedia Views
The chart below shows pageviews of the English Wikipedia article Design of experiments, from July 2015 to November 2021.[354]

Meta information on the timeline
How the timeline was built
The initial version of the timeline was written by Sebastian Sanchez.
Funding information for this timeline is available.
Feedback and comments
Feedback for the timeline can be provided at the following places:
- FIXME
What the timeline is still missing
- Vipul: "will this timeline eventually talk of things like double-blinding✔ , triple-blinding✔ , placebos✔ , RCTs✔ , etc., right? You have blinding but I guess the rest are variants on the idea".✔
- Vipul: "Cover "Statistical significance"✔ , "p-values" and preregistration."✔
- Glossary of experimental design
- Category:Design of experiments
- Design of experiments (check See also list)
Timeline update strategy
See also
External links
References
- ↑ 1.0 1.1 "Clinical trials—from ancient Babylon to today". Main Line Health. 17 May 2021. Retrieved 31 March 2026.
- ↑ Kleisiaris, Christos F.; Sfakianakis, Chrisanthos; Papathanasiou, Ioanna V. (March 15, 2014). "Health care practices in ancient Greece: The Hippocratic ideal". Journal of Medical Ethics and History of Medicine. 7. PMC 4263393. PMID 25512827. Retrieved April 9, 2025.
- ↑ "Baconian method". Encyclopaedia Britannica. Retrieved April 9, 2025.
- ↑ Lienhard, John H. (January 4, 1989). "Galileo's Experiment". The Engines of Our Ingenuity. University of Houston. Retrieved April 9, 2025.
- ↑ "Robert Boyle". Internet Encyclopedia of Philosophy. Retrieved April 9, 2025.
- ↑ Ore, Øystein (May 1960). "Pascal and the Invention of Probability Theory" (PDF). The American Mathematical Monthly. 67 (5). Mathematical Association of America: 409–419. JSTOR 2309286. Retrieved April 9, 2025.
- ↑ "Fermat and Pascal on Probability" (PDF). University of York. Retrieved April 9, 2025.
- ↑ Polasek, Wolfgang (August 2000). "The Bernoullis and the Origin of Probability Theory: Looking back after 300 Years". Resonance – Journal of Science Education. 5 (8): 26–42. Retrieved April 9, 2025.
- ↑ Cline, Douglas (August 9, 2020). "Age of Enlightenment". Physics LibreTexts. University of Rochester. Retrieved April 9, 2025.
- ↑ "The 'father of modern statistics' honoured". BBC News. September 9, 2016. Retrieved April 9, 2025.
- ↑ Schulz, Kathryn (August 14, 2023). "How Carl Linnaeus Set Out to Label All of Life". The New Yorker. Retrieved April 9, 2025.
- ↑ Schwarz, K. A., & Pfister, R.: Scientific psychology in the 18th century: a historical rediscovery. In: Perspectives on Psychological Science, Nr. 11, p. 399-407.
- ↑ "Statement on R A Fisher". Rothamsted Research. June 2020. Retrieved April 9, 2025.
- ↑ 14.00 14.01 14.02 14.03 14.04 14.05 14.06 14.07 14.08 14.09 14.10 14.11 14.12 "1.1 - A Quick History of the Design of Experiments (DOE) | STAT 503". PennState: Statistics Online Courses. Retrieved 11 May 2021.
- ↑ Preece, D. A. (December 1990). "R. A. Fisher and Experimental Design: A Review". Biometrics. 46 (4). International Biometric Society: 925–935. doi:10.2307/2532438. JSTOR 2532438. Retrieved April 9, 2025.
- ↑ Biau, David Jean; Jolles, Brigitte M.; Porcher, Raphaël (March 2010). "P Value and the Theory of Hypothesis Testing: An Explanation for New Researchers". Clinical Orthopaedics and Related Research. 468 (3): 885–892. doi:10.1007/s11999-009-1164-4. PMC 2816758. PMID 19921345. Retrieved April 9, 2025.
- ↑ Kramer, Lloyd; Maza, Sarah (23 June 2006). A Companion to Western Historical Thought. Wiley. ISBN 978-1-4051-4961-7.
Shortly after the start of the Cold War [...] double-blind reviews became the norm for conducting scientific medical research, as well as the means by which peers evaluated scholarship, both in science and in history.
- ↑ 18.0 18.1 18.2 18.3 18.4 18.5 18.6 18.7 "Chapter 2 A Brief History of Experimental Design". JABSTB: Statistical Design and Analysis of Experiments with R. Retrieved 2026-04-07.
- ↑ 19.00 19.01 19.02 19.03 19.04 19.05 19.06 19.07 19.08 19.09 19.10 19.11 19.12 19.13 19.14 "Experimental Design". Encyclopedia.com. Gale. Retrieved 2026-04-07.
- ↑ 20.0 20.1 20.2 20.3 20.4 20.5 20.6 20.7 20.8 "A Brief History of Statistical Design in Experiments (DAE)". Studocu. Retrieved 21 August 2026.
- ↑ 21.0 21.1 21.2 21.3 21.4 21.5 21.6 "Designing of Experiments". StudyAndScore. Retrieved 9 April 2026.
- ↑ 22.00 22.01 22.02 22.03 22.04 22.05 22.06 22.07 22.08 22.09 22.10 22.11 Chandramouli, R. "100 years of DoE – from Fisher to Jones". LinkedIn. Retrieved 2026-04-07.
- ↑ "Textbooks and other publications on controlled clinical trials". PubMed Central. Retrieved 3 April 2026.
- ↑ Davis, Rahul; John, Pretesh (7 March 2018). "Application of Taguchi-Based Design of Experiments for Industrial Chemical Processes". IntechOpen. Statistical Approaches With Emphasis on Design of Experiments Applied to Chemical Processes.
- ↑ "Design of experiments". Wikipedia. Retrieved 3 April 2026.
- ↑ "Randomized controlled trial - History". Wikipedia. Retrieved 3 April 2026.
- ↑ "Randomized controlled trial - applications". Wikipedia. Retrieved 3 April 2026.
- ↑ Lancaster, Laura. "Disallowed Combinations and Operating Region Optimization for CQAs with the JMP 17 Profiler". JMP User Community, Discovery Summit Europe 2023.
- ↑ "Randomized controlled trial - applications". Wikipedia. Retrieved 3 April 2026.
- ↑ Gerhard von Rad, Old Testament Theology, Vol. 2 (Louisville: Westminster John Knox, 1965), 307–308.
- ↑ David C. Lindberg, The Beginnings of Western Science (Chicago: University of Chicago Press, 2007), 14–15.
- ↑ James C. VanderKam, An Introduction to Early Judaism (Grand Rapids: Eerdmans, 2001), 124.
- ↑ Gerhard von Rad, Old Testament Theology, Vol. 2 (Louisville: Westminster John Knox, 1965), 307–308.
- ↑ David C. Lindberg, The Beginnings of Western Science (Chicago: University of Chicago Press, 2007), 14–15.
- ↑ James C. VanderKam, An Introduction to Early Judaism (Grand Rapids: Eerdmans, 2001), 124.
- ↑ Hayashi, Takao (2008). "Magic Squares in Indian Mathematics". Encyclopaedia of the History of Science, Technology, and Medicine in Non-Western Cultures (2nd ed.). Springer. pp. 1252–1259. doi:10.1007/978-1-4020-4425-0_9778.
- ↑ El-Bizri, Nader (2005). "A Philosophical Perspective on Alhazen's Optics". Arabic Sciences and Philosophy. 15 (2): 189–218. doi:10.1017/S0957423905000172.
- ↑ Aligabi, Zahra (2020). "Reflections on Avicenna's impact on medicine: his reach beyond the Middle East". Journal of Community Hospital Internal Medicine Perspectives. 10 (4): 310–312. doi:10.1080/20009666.2020.1774301. Retrieved 31 March 2026.
- ↑ Pearce, J. M. S. (29 September 2022). "Francis Bacon's natural philosophy and medicine". Hektoen International: A Journal of Medical Humanities. Retrieved 31 March 2026.
- ↑ Donaldson, I.M.L. (September 2016). "Van Helmont's Proposal for a Randomised Comparison of Treating Fevers with or without Bloodletting and Purging". Journal of the Royal College of Physicians of Edinburgh. 46 (3): 206–213. doi:10.4997/jrcpe.2016.313. PMID 27959358. Retrieved 29 August 2026.
- ↑ "Ole Rømer Profile: First to Measure the Speed of Light". American Museum of Natural History. Retrieved 13 August 2026.
- ↑ Boyle, Robert (1683). New Experiments and Observations Touching Cold, or, An Experimental History of Cold, Begun. London: Richard Davis. Retrieved 31 March 2026.
- ↑ Colbourn, Charles J.; Dinitz, Jeffrey H. Handbook of Combinatorial Designs (2nd ed.). CRC Press. p. 12. ISBN 9781420010541. Retrieved 28 March 2017.
- ↑ "Choi Seok-jeong (1646–1715)". MacTutor History of Mathematics Archive. University of St Andrews. Retrieved 31 March 2026.
- ↑ John Arbuthnot (1710). "An argument for Divine Providence, taken from the constant regularity observed in the births of both sexes". Philosophical Transactions of the Royal Society of London. 27 (325–336): 186–190. doi:10.1098/rstl.1710.0011.
- ↑ 46.0 46.1 Dunn, Peter M. (January 1, 1997). "James Lind (1716-94) of Edinburgh and the treatment of scurvy". Archives of Disease in Childhood: Fetal and Neonatal Edition. 76 (1): F64–5. doi:10.1136/fn.76.1.F64. PMC 1720613. PMID 9059193.
- ↑ 47.0 47.1 47.2 Kerr, C.E.; Milne, I.; Kaptchuk, T.J. (February 2008). "William Cullen and a missing mind-body link in the early history of placebos". Journal of the Royal Society of Medicine. 101 (2): 89–92. doi:10.1258/jrsm.2007.071005. PMC 2254457. PMID 18299629. Retrieved 29 August 2026.
- ↑ Laplace, P. (1778). "Mémoire sur les probabilités". Mémoires de l'Académie Royale des Sciences de Paris: 227–332.
- ↑ Euler, Leonhard (1782). "Recherches sur une nouvelle espèce de quarrés magiques" [Investigations into a new type of magic squares]. Verhandelingen Uitgegeven Door Het Zeeuwsch Genootschap der Wetenschappen te Vlissingen (in French). 9: 85–239.
{{cite journal}}: CS1 maint: unrecognized language (link) - ↑ Donaldson, I. M. L. (December 2005). "Mesmer's 1780 Proposal for a Controlled Trial to Test his Method of Treatment Using 'Animal Magnetism'". Journal of the Royal Society of Medicine. 98 (12): 572–575. doi:10.1177/014107680509801226.
- ↑ "Kent Academic Repository" (PDF). kar.kent.ac.uk. Retrieved 23 October 2021.
- ↑ Donaldson, I. M. L. (April 2017). "Antoine de Lavoisier's Role in Designing a Single-Blind Trial to Assess whether 'Animal Magnetism' Exists". Journal of the Royal Society of Medicine. 110 (4): 163–167. doi:10.1177/0141076817699740.
- ↑ Gould, Stephen J. (1989). "The Chain of Reason vs. The Chain of Thumbs". Natural History. 98 (7): 12–21.
- ↑ 54.0 54.1 Aronson, Jeff (13 March 1999). "Please, please me". BMJ. 318 (7185): 716. doi:10.1136/bmj.318.7185.716. PMC 1115150. PMID 10074020. Retrieved 29 August 2026.
- ↑ Riedel, Stefan (January 2005). "Edward Jenner and the history of smallpox and vaccination". Proceedings (Baylor University Medical Center). pp. 21–25. Retrieved 21 August 2026.
- ↑ "Carl Friedrich Gauss & Adrien-Marie Legendre Discover the Method of Least Squares". History of Information. Jeremy M. Norman. Retrieved 31 March 2026.
- ↑ Booth, Christopher (August 2005). "The rod of Aesculapios: John Haygarth (1740-1827) and Perkins' metallic tractors". Journal of Medical Biography. pp. 155–161. Retrieved 21 August 2026.
- ↑ "Polynomial regression". frontend. Retrieved 18 March 2022.
- ↑ Fétis, François-Joseph (1868). Biographie Universelle des Musiciens et Bibliographie Générale de la Musique, Tome 1 (Second ed.). Paris: Firmin Didot Frères, Fils, et Cie. p. 249. Retrieved 2011-07-21.
- ↑ Dubourg, George (1852). The Violin: Some Account of That Leading Instrument and its Most Eminent Professors... (Fourth ed.). London: Robert Cocks and Co. pp. 356–357. Retrieved 2011-07-21.
- ↑ Stigler (1986, pp 154–155)
- ↑ Stolberg, M. (December 2006). "Inventing the randomized double-blind trial: the Nuremberg salt test of 1835". Journal of the Royal Society of Medicine. 99 (12): 642–643. doi:10.1258/jrsm.99.12.642. PMC 1676327. PMID 17139070.
- ↑ Faerstein, Eduardo; Winkelstein Jr., Warren (September 2012). "Adolphe Quetelet: Statistician and More". Epidemiology. 23 (5): 762–763. doi:10.1097/EDE.0b013e318261c86f. Retrieved 31 March 2026.
- ↑ Cooper, Max; Middleton, Jo; Cooper, Sarah (2025). "Cold-water, Sulphur and 'the itch': James Henry's principles for conducting controlled trials (1843)". Irish Journal of Medical Science. 194 (6): 2303–2305. doi:10.1007/s11845-025-04027-x. PMC 12769593. PMID 40824558.
{{cite journal}}: Check|pmc=value (help); Check|pmid=value (help) - ↑ Lindner & Rodger 1997, pg.3
- ↑ Kirkman, Thomas P. (1847), "On a Problem in Combinations", The Cambridge and Dublin Mathematical Journal, II: 191–204
- ↑ J J O'Connor; E F Robertson (December 1996). "Thomas Penyngton Kirkman". MacTutor History of Mathematics Archive. University of St Andrews. Retrieved 21 August 2026.
- ↑ Steiner, J. (1853), "Combinatorische Aufgabe", Journal für die reine und angewandte Mathematik, 1853 (45): 181–182, doi:10.1515/crll.1853.45.181
- ↑ Tulchinsky, Theodore H. (30 March 2018). "John Snow, Cholera, the Broad Street Pump; Waterborne Diseases Then and Now". Case Studies in Public Health. Elsevier. pp. 77–99. Retrieved 21 August 2026.
- ↑ 70.0 70.1 "Experimental Psychology: History, Features, and Method". Psychologs Magazine. Retrieved April 9, 2025.
- ↑ Pasteur, Louis (1861). Sur les corpuscules organisés qui existent dans l'atmosphère: Examen de la doctrine des générations spontanées (in français). Paris: Ch. Lahure et Cie. Retrieved 31 March 2026.
- ↑ Cavaillon, Jean-Marc; Legout, Sandra (2022). "Louis Pasteur: Between Myth and Reality". Biomolecules. 12 (4): 596. doi:10.3390/biom12040596. PMC 9027159. PMID 35454184. Retrieved 31 March 2026.
{{cite journal}}: CS1 maint: unflagged free DOI (link) - ↑ Bernard, Claude (2008) [1865]. Introduction à l'étude de la médecine expérimentale. Champs. Paris: Flammarion. ISBN 978-2-08-121793-5.
- ↑ Peirce, C. S. (August 1967). "Note on the Theory of the Economy of Research". Operations Research. 15 (4): 643–648. doi:10.1287/opre.15.4.643.
- ↑ Robert Burch (2001). "Charles Sanders Peirce". Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University. Retrieved 6 April 2026.
- ↑ Peirce, Charles Sanders (1877–1878). "Illustrations of the Logic of Science". Popular Science Monthly. Retrieved 13 August 2026.
- ↑ 77.0 77.1 77.2 Kaptchuk, Ted J. (2011). "A brief history of the evolution of methods to control observer biases in tests of treatments". James Lind Library, JLL Bulletin: Commentaries on the history of treatment evaluation. Retrieved 29 August 2026.
- ↑ Kaptchuk, Ted J. (2004). "Early use of blind assessment in a homeopathic scientific experiment". James Lind Library. Retrieved 29 August 2026.
- ↑ Peirce, C. S. (1882), "Introductory Lecture on the Study of Logic" delivered September 1882, published in Johns Hopkins University Circulars, v. 2, n. 19, pp. 11–12, November 1882, see p. 11, Google Books Eprint. Reprinted in Collected Papers v. 7, paragraphs 59–76, see 59, 63, Writings of Charles S. Peirce v. 4, pp. 378–82, see 378, 379, and The Essential Peirce v. 1, pp. 210–14, see 210–1, also lower down on 211.
- ↑ O'Rourke, M.F. (1992). "Frederick Akbar Mahomed". Hypertension. 19 (2): 212–217. doi:10.1161/01.HYP.19.2.212.
- ↑ Stigler (1986, pp 314–315)
- ↑ "Role of the Michelson-Morley experiments in making determinations about competing theories". Archived from the original on 2012-11-07. Retrieved 2003-07-17.
- ↑ Charles Sanders Peirce and Joseph Jastrow (1885). "On Small Differences in Sensation". Memoirs of the National Academy of Sciences. 3: 73–83. http://psychclassics.yorku.ca/Peirce/small-diffs.htm
- ↑ Binet, Alfred (15 October 1894). "La psychologie de la prestidigitation". Alfred Binet (1857-1911) — Works Archive. Serge Nicolas, Paris Descartes University. Retrieved 21 August 2026.
- ↑ Binet, Alfred (1896). "Psychology of Prestidigitation". Annual Report of the Board of Regents of the Smithsonian Institution showing the Operations, Expenditures, and Condition of the Institution to July 1894, pp. 555–571. Government Printing Office.
{{cite web}}: Missing or empty|url=(help) - ↑ Stroebe, W. (2012). The truth about Triplett (1898), but nobody seems to care. Perspectives on Psychological Science, 7, 54-57.
- ↑ Pearson, Karl (1900). "On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling" (PDF). Philosophical Magazine. Series 5. 50 (302): 157–175. doi:10.1080/14786440009463897.
- ↑ Nahm, Francis Sahngun (2017). "What the P values really tell us". The Korean Journal of Pain. 30 (4): 241. doi:10.3344/kjp.2017.30.4.241.
- ↑ Nahm, Francis Sahngun (October 2017). "What the P values really tell us". The Korean Journal of Pain. 30 (4): 241–242. doi:10.3344/kjp.2017.30.4.241. ISSN 2005-9159.
- ↑ Raúl Ibáñez (28 January 2015). "Los cuadrados greco-latinos de Leonhard Euler". Cuaderno de Cultura Científica. Cátedra de Cultura Científica, UPV/EHU. Retrieved 21 August 2026.
- ↑ Lange, Katie (5 February 2021). "Walter Reed: Get to Know the Man Behind the Medical Center". U.S. Department of War. Retrieved 21 August 2026.
- ↑ Newman, David H., M.D. (2008). Hippocrates' shadow : secrets from the house of medicine (1st Scribner hardcover ed.). New York, NY: Scribner. ISBN 978-1-4165-5153-9.
{{cite book}}: CS1 maint: multiple names: authors list (link) - ↑ 93.0 93.1 Mulaik, Stanley A. (2010). Foundations of Factor Analysis (2nd ed.). Boca Raton, Florida: CRC Press. p. 6. ISBN 978-1-4200-9961-4.
- ↑ Spearman, Charles (1904). "General intelligence, objectively determined and measured". American Journal of Psychology. 15 (2): 201–293. doi:10.2307/1412107. JSTOR 1412107.
- ↑ Samhita, Laasya; Gross, Hans J. (1 November 2013). "The "Clever Hans Phenomenon" revisited". Communicative & Integrative Biology. 6 (6) e27122. doi:10.4161/cib.27122. PMC 3921203. PMID 24563716.
- ↑ Rivers WH, Webber HN (August 1907). "The action of caffeine on the capacity for muscular work". The Journal of Physiology. 36 (1): 33–47. doi:10.1113/jphysiol.1907.sp001215. PMC 1533733. PMID 16992882.
- ↑ Gimhan, Vishal (6 March 2025). "Understanding the t-Distribution: A Guide for Small Sample Analysis". Medium. Retrieved 31 March 2026.
- ↑ Edgeworth, F.Y. (June 1908). "On the Probable Errors of Frequency-Constants". Journal of the Royal Statistical Society. 71 (2): 381–397. JSTOR 2339461.
- ↑ The Correlation Between Relatives on the Supposition of Mendelian Inheritance. Ronald A. Fisher. Philosophical Transactions of the Royal Society of Edinburgh. 1918. (volume 52, pages 399–433)
- ↑ Smith, Kirstine (1918). "On the Standard Deviations of Adjusted and Interpolated Values of an Observed Polynomial Function and its Constants and the Guidance they give Towards a Proper Choice of the Distribution of Observations". Biometrika. 12 (1–2). Oxford University Press: 1–85. doi:10.2307/2331929. Retrieved 31 March 2026.
- ↑ 101.0 101.1 "Experimental Design | Encyclopedia.com". www.encyclopedia.com. Retrieved 5 April 2021.
- ↑ Graves, T.C. (1920). "Commentary on a case of Hystero-epilepsy with delayed puberty". The Lancet. 196 (5075): 1135. doi:10.1016/S0140-6736(01)00108-8.
{{cite journal}}:|access-date=requires|url=(help) - ↑ On the "Probable Error" of a Coefficient of Correlation Deduced from a Small Sample. Ronald A. Fisher. Metron, 1: 3–32 (1921)
- ↑ Fisher, R.A. (1922). "On the mathematical foundations of theoretical statistics". Philosophical Transactions of the Royal Society of London, Series A. 222 (594–604): 309–368. doi:10.1098/rsta.1922.0009.
- ↑ Scheffé (1959, p 291, "Randomization models were first formulated by Neyman (1923) for the completely randomized design, by Neyman (1935) for randomized blocks, by Welch (1937) and Pitman (1937) for the Latin square under a certain null hypothesis, and by Kempthorne (1952, 1955) and Wilk (1955) for many other designs.")
- ↑ Fisher, Ronald A. (1921). "Studies in Crop Variation. I. An Examination of the Yield of Dressed Grain from Broadbalk". Journal of Agricultural Science. 11 (2): 107–135. doi:10.1017/S0021859600003750.
- ↑ Fisher, Ronald A. (1923). "Studies in Crop Variation. II. The Manurial Response of Different Potato Varieties". Journal of Agricultural Science. 13 (3): 311–320. doi:10.1017/S0021859600003592.
- ↑ Gosnell, Harold F. (1926). "An Experiment in the Stimulation of Voting". American Political Science Review. 20 (4): 869–874. doi:10.1017/S0003055400110524.
- ↑ Cumming, Geoff (2011). "From null hypothesis significance to testing effect sizes". Understanding The New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis. Multivariate Applications Series. East Sussex, United Kingdom: Routledge. pp. 21–52. ISBN 978-0-415-87968-2.
- ↑ Fisher, Ronald A. (1925). Statistical Methods for Research Workers. Edinburgh, UK: Oliver and Boyd. pp. 43. ISBN 978-0-050-02170-5.
{{cite book}}: ISBN / Date incompatibility (help) - ↑ Lehmann, Erich L. (2011). Fisher, Neyman, and the creation of classical statistics. New York, NY: Springer Science+Business Media, LLC. p. 15. ISBN 978-1-4419-9500-1.
- ↑ Conniffe, Denis (1990–1991). "R. A. Fisher and the development of statistics—a view in his centenary year". Journal of the Statistical and Social Inquiry Society of Ireland. Vol. XXVI, no. 3. Dublin: Statistical and Social Inquiry Society of Ireland. p. 87. hdl:2262/2764. ISSN 0081-4776.
- ↑ Savage, Leonard J. (1976). "On Rereading R. A. Fisher". Annals of Statistics. 4 (3): 441–500. doi:10.1214/aos/1176343456.
- ↑ Kopf, Dan. "An error made in 1925 led to a crisis in modern science—now researchers are joining to fix it". Quartz. Retrieved 13 March 2021.
- ↑ Box, Joan Fisher (February 1980). "R. A. Fisher and the Design of Experiments, 1922-1926". The American Statistician. 34 (1): 1. doi:10.2307/2682986.
- ↑ Press, David J.; Pharoah, Paul (July 2010). "Risk factors for breast cancer: a reanalysis of two case-control studies from 1926 and 1931". Epidemiology. pp. 566–572. Retrieved 21 August 2026.
- ↑ Fisher, Ronald (1926). "The Arrangement of Field Experiments". Journal of the Ministry of Agriculture of Great Britain. 33. London: Ministry of Agriculture and Fisheries: 503–513.
- ↑ Bernstein, S. N. (1926). "Sur l'extension du théorème limite du calcul des probabilités aux sommes de quantités dépendantes". Mathematische Annalen. 97: 1–59.
- ↑ 119.0 119.1 119.2 119.3 Walach, Harald (2015). "Placebo Studies (Double-Blind Studies)". International Encyclopedia of the Social & Behavioral Sciences (2nd ed.). Elsevier. doi:10.1016/B978-0-08-097086-8.03124-2. Retrieved 29 August 2026.
- ↑ Thurstone, Louis (1931). "Multiple factor analysis". Psychological Review. 38 (5): 406–427. doi:10.1037/h0069792.
- ↑ Thurstone, L.L. (1935). The Vectors of Mind: Multiple-Factor Analysis for the Isolation of Primary Traits. Chicago: University of Chicago Press.
- ↑ Neyman, J; Pearson, E. S. (1 January 1933). "On the Problem of the most Efficient Tests of Statistical Hypotheses". Philosophical Transactions of the Royal Society A. 231 (694–706): 289–337. Bibcode:1933RSPTA.231..289N. doi:10.1098/rsta.1933.0009.
- ↑ Paley, R. E. A. C.; Lin, C. Devon; Stufken, John (1933). "On orthogonal matrices". Orthogonal Arrays: A Review (arXiv:2505.15032), citing Paley (1933), Journal of Mathematics and Physics 12(1-4), 311–320. Retrieved 21 August 2026.
- ↑ Evans, William; Hoyle, Clifford (1933). "The comparative value of drugs used in the continuous treatment of angina pectoris". Quarterly Journal of Medicine. Retrieved 29 August 2026.
- ↑ Fienberg, Stephen E.; Tanur, Judith M. (2001). "Jerzy Neyman". Statisticians of the Centuries. Springer. pp. 444–448. Retrieved 21 August 2026.
- ↑ Fisher, R.A. (1935). The Design of Experiments. Oliver and Boyd. pp. 114–145.
- ↑ Box, JF (February 1980). "R. A. Fisher and the Design of Experiments, 1922–1926". The American Statistician. 34 (1): 1–7. doi:10.2307/2682986. JSTOR 2682986.
- ↑ Yates, F (June 1964). "Sir Ronald Fisher and the Design of Experiments". Biometrics. 20 (2): 307–321. doi:10.2307/2528399. JSTOR 2528399.
- ↑ Stanley, Julian C. (1966). "The Influence of Fisher's "The Design of Experiments" on Educational Research Thirty Years Later". American Educational Research Journal. 3 (3): 223–229. doi:10.3102/00028312003003223. JSTOR 1161806.
- ↑ Fisher, Ronald A. (1971) [1935]. The Design of Experiments (9th ed.). Macmillan. ISBN 0-02-844690-9.
- ↑ Box, Joan Fisher (1978). R.A. Fisher, The Life of a Scientist. New York: Wiley. p. 134. ISBN 0-471-09300-9.
- ↑ "Earliest Known Uses of Some of the Words of Mathematics (F)". jeff560.tripod.com. Retrieved 13 August 2026.
- ↑ Crawford, Meredith P. (1937). The Coöperative Solving of Problems by Young Chimpanzees. Johns Hopkins Press.
- ↑ Gold, Harry; Kwit, Nathaniel T.; Otto, Harold (1937). "The Xanthines (Theobromine and Aminophyllin) in the treatment of cardiac pain". Journal of the American Medical Association. 108 (26): 2173. doi:10.1001/jama.1937.02780260001001.
{{cite journal}}:|access-date=requires|url=(help) - ↑ Yates, F. (1937). The Design and Analysis of Factorial Experiments. Harpenden, England: Commonwealth Bureau of Soils. Technical Communication 35, pp. 66–67.
- ↑ Bose, R. C.; Nair, K. R. (1939), "Partially balanced incomplete block designs", Sankhyā, 4: 337–372
- ↑ Fisher, R.A. (1940), "An examination of the different possible solutions of a problem in incomplete blocks", Annals of Eugenics, 10: 52–75, doi:10.1111/j.1469-1809.1940.tb02237.x, hdl:2440/15239
- ↑ Kishen, K. (1942), "On latin and hyper-graeco cubes and hypercubes", Current Science, 11: 98–99
- ↑ Dorfman, Robert (December 1943). "The Detection of Defective Members of Large Populations". JSTOR. Institute of Mathematical Statistics. Retrieved 21 August 2026.
- ↑ "Clinical trial of patulin in the common cold. 1944". International Journal of Epidemiology. April 2004.
- ↑ Chalmers, I.; Clarke, M. (April 2004). "Commentary: the 1944 patulin trial: the first properly controlled multicentre trial conducted under the aegis of the British Medical Research Council". International Journal of Epidemiology.
- ↑ Wald, Abraham (1945). "Sequential Tests of Statistical Hypotheses". Annals of Mathematical Statistics.
- ↑ Finney, D. J. (1945). "The fractional replication of factorial arrangements". Annals of Eugenics.
- ↑ National Research Council (1995). Statistical Methods for Testing and Evaluating Defense Systems: Interim Report. Washington, DC: The National Academies Press. doi:10.17226/9074.
- ↑ Jellinek, E. M. "Clinical Tests on Comparative Effectiveness of Analgesic Drugs", Biometrics Bulletin, Vol.2, No.5, (October 1946), pp.87–91.
- ↑ "5.3.3.5. Plackett-Burman designs". www.itl.nist.gov. Retrieved 22 July 2023.
- ↑ Rao, C.R. (1946). "Hypercubes of strength "d" leading to confounded designs in factorial experiments". Bulletin of the Calcutta Mathematical Society, vol. 38, pp. 67–78. Indian Academy of Sciences repository.
- ↑ Rao, C.R. (1947). "Factorial experiments derivable from combinatorial arrangements of arrays". Supplement to the Journal of the Royal Statistical Society.
- ↑ Metcalfe, N.H. (2011). "Sir Geoffrey Marshall (1887-1982): respiratory physician, catalyst for anaesthesia development, doctor to both Prime Minister and King, and World War I Barge Commander". Journal of Medical Biography. 19 (1): 10–14. doi:10.1258/jmb.2010.010019. PMID 21350072.
- ↑ "The MRC randomized trial of streptomycin and its legacy". PubMed Central. Retrieved 3 April 2026.
- ↑ "Nuremberg Code". The Doctor's Trial: The Medical Case of the Subsequent Nuremberg Proceedings. United States Holocaust Memorial Museum Online Exhibitions. Retrieved 13 February 2019.
- ↑ Shuster, Evelyne (1997). "Fifty Years Later: The Significance of the Nuremberg Code". New England Journal of Medicine. 337 (20): 1436–1440. doi:10.1056/NEJM199711133372006. PMID 9358142.
- ↑ Anscombe, F. J. (1948). "The Validity of Comparative Experiments". Journal of the Royal Statistical Society. Series A (General).
- ↑ Healy, M. J. R. (1995). "Frank Yates, 1902-1994: The Work of a Statistician". International Statistical Review / Revue Internationale de Statistique.
- ↑ Grundy, P. M.; Healy, M. J. R. (1950). "Restricted Randomization and Quasi-Latin Squares". Journal of the Royal Statistical Society. Series B (Methodological).
- ↑ Navarro, Mario; Siegel, Jason T. (2018). "Solomon Four-Group Design". SAGE Publications. Retrieved 22 November 2019.
- ↑ Moller-Wong, Cheryl Lynn (1988). "The Taguchi Methods of Quality Control Examined: With Reference to Sundstrand-Sauer". Iowa State University. Retrieved 2 March 2026.
- ↑ Arrow, Kenneth J.; Blackwell, David; Girshick, M.A. (1949). "Bayes and minimax solutions of sequential decision problems". Econometrica.
- ↑ Bruck, R.H.; Ryser, H.J. (1949). "The nonexistence of certain finite projective planes". Canadian Journal of Mathematics.
- ↑ Agresti, Alan (2023). "A historical overview of textbook presentations of statistical science" (PDF). Scandinavian Journal of Statistics. pp. 1641–1666. Retrieved 21 August 2026.
- ↑ Grundy, P.M.; Healy, M.J.R. (1950). "Restricted randomization and quasi-Latin squares". Journal of the Royal Statistical Society, Series B.
- ↑ Bush, K. A. (1950). "Orthogonal arrays". University of North Carolina.
{{cite web}}: Missing or empty|url=(help) - ↑ Doll, Richard; Hill, Austin Bradford (1950). "Smoking and carcinoma of the lung; preliminary report". British Medical Journal.
- ↑ Chowla, S.; Ryser, H.J. (1950). "Combinatorial problems". Canadian Journal of Mathematics.
- ↑ Draper, Norman R. (1992). "Introduction to Box and Wilson (1951) On the Experimental Attainment of Optimum Conditions". Breakthroughs in Statistics: Methodology and Distribution. Springer. pp. 267–269. doi:10.1007/978-1-4612-4380-9_22.
- ↑ Martins, Joaquim R. R. A.; Ning, Andrew (January 2022). "A Short History of Optimization". Engineering Design Optimization. Cambridge University Press. doi:10.1017/9781108980647.
- ↑ Robbins, Herbert (1952). "Some aspects of the sequential design of experiments". Bulletin of the American Mathematical Society.
- ↑ Horvitz, D. G.; Thompson, D. J. (1952). "A generalization of sampling without replacement from a finite universe". Journal of the American Statistical Association.
- ↑ Bose, R. C.; Shimamoto, T. (June 1952). "Classification and Analysis of Partially Balanced Incomplete Block Designs with Two Associate Classes". Journal of the American Statistical Association.
- ↑ Hróbjartsson A, Gøtzsche PC (May 2001). "Is the placebo powerless? An analysis of clinical trials comparing placebo with no treatment". The New England Journal of Medicine. 344 (21): 1594–602. doi:10.1056/NEJM200105243442106. PMID 11372012.
- ↑ Wilk, M.B. (1955). "The Randomization Analysis of a Generalized Randomized Block Design". Biometrika.
- ↑ Lindley, D. V. (1956). "On a measure of information provided by an experiment". Annals of Mathematical Statistics.
- ↑ Doll, Richard; Hill, Austin Bradford (1956). "Lung cancer and other causes of death in relation to smoking; a second report on the mortality of British doctors". British Medical Journal.
- ↑ Kish, L. (1959). "Some statistical problems in research design". American Sociological Review.
- ↑ Pierson, Raymond H.; Fay, Edward A. (December 1959). "Guidelines for Interlaboratory Testing Programs". Analytical Chemistry.
- ↑ Bose, R. C.; Mesner, D. M. (1959). "On linear associative algebras corresponding to association schemes of partially balanced designs". Annals of Mathematical Statistics.
- ↑ Russell, W.M.S.; Burch, R.L. (1959). The Principles of Humane Experimental Technique. London: Methuen. ISBN 0-900767-78-2.
{{cite book}}: ISBN / Date incompatibility (help) - ↑ Samuel, Arthur L. (31 July 1959). "Some Studies in Machine Learning Using the Game of Checkers". IBM Journal of Research and Development. pp. 210–229. Retrieved 21 August 2026.
- ↑ "Genichi Taguchi". asq.org. American Society for Quality. Retrieved 31 March 2026.
- ↑ Box, George E. P.; Behnken, Donald (1960). "Some new three level designs for the study of quantitative variables". Technometrics.
- ↑ Ranade, Shruti Sunil; Thiagarajan, Padma (November 2017). "Selection of a design for response surface". IOP Conference Series: Materials Science and Engineering.
- ↑ Raghavarao, Damaraju (1960). "Some Aspects of Weighing Designs". Annals of Mathematical Statistics.
- ↑ Wason, Peter C. (1960). "On the failure to eliminate hypotheses in a conceptual task". Quarterly Journal of Experimental Psychology.
- ↑ Thistlethwaite, D.; Campbell, D. (1960). "Regression-Discontinuity Analysis: An alternative to the ex post facto experiment". Journal of Educational Psychology.
- ↑ Box, G. E. P.; Hunter, J. S. (August 1961). "The 2^(k-p) Fractional Factorial Designs, Part I". Technometrics. pp. 311–351.
- ↑ Connell, Joseph H. (October 1961). "The Influence of Interspecific Competition and Other Factors on the Distribution of the Barnacle Chthamalus Stellatus". Ecology. Wiley. pp. 710–723. Retrieved 21 August 2026.
- ↑ "Definition of NOCEBO". www.merriam-webster.com. Retrieved 5 March 2022.
- ↑ Kennedy, 1961
- ↑ Harman, WW; McKim, RH; Mogar, RE; Fadiman, J; Stolaroff, MJ (August 1966). "Psychedelic agents in creative problem-solving: a pilot study". Psychological Reports.
- ↑ Kôno, Kazumasa (1962). "Optimum designs for quadratic regression on k-cube". Memoirs of the Faculty of Science. Kyushu University. Series A. Mathematics.
- ↑ Smith, Vernon L. (1962). "An Experimental Study of Competitive Market Behavior". Chapman University Digital Commons. Chapman University. Retrieved 31 March 2026.
- ↑ Stankova, Tatiana (30 June 2020). "Application of Nelder wheel experimental design in forestry research". Silva Balcanica.
- ↑ 193.0 193.1 Heppner, Puncky Paul; Wampold, Bruce E.; Owen, Jesse; Wang, Kenneth T. (21 August 2015). Research Design in Counseling. Cengage Learning. ISBN 978-1-305-46501-5.
- ↑ Boren, J. J. (1963). "Repeated acquisition of new behavioral chains". American Psychologist, 18, p. 421.
- ↑ "Declaration of Helsinki (1964)". CIRP.org, reprinted from British Medical Journal. 7 December 1996. pp. 1448–1449. Retrieved 21 August 2026.
- ↑ Kish, Leslie (1965). Survey Sampling. New York: John Wiley & Sons, Inc. ISBN 0-471-10949-5.
- ↑ Selvin, H.C.; Stuart, A. (1966). "Data-Dredging Procedures in Survey Analysis". The American Statistician.
- ↑ Rosenthal, R. (1966). Experimenter Effects in Behavioral Research. New York: Appleton-Century-Crofts.
- ↑ Myers, Raymond H. (1971). Response Surface Methodology. Boston: Allyn and Bacon.
- ↑ Chernoff, H. (1972) Sequential Analysis and Optimal Design, SIAM Monograph.
- ↑ Delsarte, P. (1973). "An Algebraic Approach to the Association Schemes of Coding Theory". Philips Research Reports, Supplement No. 10.
- ↑ Rubin, Donald B. (1973). "Matching to Remove Bias in Observational Studies". Biometrics.
- ↑ Wike, Edwin L. (1973). "Water beds and sexual satisfaction: Wike's law of low odd primes (WLLOP)". Psychological Reports. pp. 192–194.
{{cite web}}: Missing or empty|url=(help) - ↑ Rubin, Donald B. (1974). "Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies". Journal of Educational Psychology.
- ↑ Denniston, R. H. F. (September 1974). "Denniston's paper, open access". Discrete Mathematics.
- ↑ Ray-Chaudhuri, Dijen K.; Wilson, Richard M. (1975). "On t-designs". Osaka Journal of Mathematics.
{{cite web}}: Missing or empty|url=(help) - ↑ Atkinson, A. C.; Fedorov, V. V. (1975). "The design of experiments for discriminating between two rival models". Biometrika.
- ↑ Pocock, Stuart J.; Simon, Richard (March 1975). "Sequential Treatment Assignment with Balancing for Prognostic Factors in the Controlled Clinical Trial". Biometrics.
- ↑ Scheirer, C. J.; Ray, W. S.; Hare, N. (June 1976). "The analysis of ranked data derived from completely randomized factorial designs". Biometrics. pp. 429–434.
- ↑ Montgomery, Douglas C. (2013). Design and Analysis of Experiments. John Wiley & Sons Incorporated. ISBN 978-1-62198-227-2.
- ↑ Sloane, Neil J. A.; Harwit, Martin (1 January 1976). "Masks for Hadamard transform optics, and weighing designs". Applied Optics.
- ↑ Miettinen, O. (1976). "Estimability and estimation in case–referent studies". American Journal of Epidemiology.
- ↑ Pocock, Stuart J. (August 1977). "Group Sequential Methods in the Design and Analysis of Clinical Trials". Biometrika.
- ↑ Eglajs, V.; Audze, P. (1977). "New approach to the design of multifactor experiments". Problems of Dynamics and Strengths (in Russian). 35. Riga: Zinatne Publishing House: 104–107.
{{cite journal}}: CS1 maint: unrecognized language (link) - ↑ "Audze–Eglais method survey and modern applications". arXiv. 2025.
- ↑ Rubin, Donald (1978). "Bayesian Inference for Causal Effects: The Role of Randomization". The Annals of Statistics.
- ↑ "Experimental Design - an overview | ScienceDirect Topics". www.sciencedirect.com. Retrieved 23 March 2021.
- ↑ Ma, Will. "Lecture 1: Four (and a Half) Proofs of the Basic Prophet Inequality" (PDF). Columbia University. Columbia Business School. Retrieved 31 March 2026.
- ↑ Richter, Felicitas; Dewey, Marc (September 2014). "Zelen Design in Randomized Controlled Clinical Trials". Radiology. 272 (3): 919–919. doi:10.1148/radiol.14140834.
- ↑ Homer, Caroline S.E. (April 2002). "Using the Zelen design in randomized controlled trials: debates and controversies". Journal of Advanced Nursing. 38 (2): 200–207. doi:10.1046/j.1365-2648.2002.02164.x.
- ↑ Zelen, Marvin (1979). "A New Design for Randomized Clinical Trials". The New England Journal of Medicine.
- ↑ Cook, Thomas D.; Campbell, Donald T. (1979). Quasi-experimentation: Design & Analysis Issues for Field Settings. Boston: Houghton-Mifflin.
- ↑ O'Brien, Peter C.; Fleming, Thomas R. (1979). "A Multiple Testing Procedure for Clinical Trials". Biometrics.
- ↑ McKay, M.D.; Beckman, R.J.; Conover, W.J. (May 1979). "A Comparison of Three Methods for Selecting Values of Input Variables in the Analysis of Output from a Computer Code". Technometrics.
- ↑ Gittins, J.C. (1979). "Bandit Processes and Dynamic Allocation Indices". Journal of the Royal Statistical Society, Series B.
- ↑ Coronary Drug Project Research Group (October 1980). "Influence of adherence to treatment and response of cholesterol on mortality in the coronary drug project". New England Journal of Medicine.
- ↑ Tukey, John W. (1980). "We Need Both Exploratory and Confirmatory". The American Statistician. 34 (1): 23–25. doi:10.2307/2682991. JSTOR 2682991.
- ↑ Karkar, Ravi; Zia, Jasmine; Vilardaga, Roger; Mishra, Sonali R; Fogarty, James; Munson, Sean A; Kientz, Julie A (1 May 2016). "A framework for self-experimentation in personalized health". Journal of the American Medical Informatics Association. 23 (3): 440–448. doi:10.1093/jamia/ocv150.
- ↑ George E.P., Box (2006). Improving Almost Anything: Ideas and Essays (Revised ed.). Hoboken, New Jersey: Wiley.
- ↑ Rubin, Donald B.; Rosenbaum, Paul R. (1983). "The Central Role of the Propensity Score in Observational Studies for Causal Effects". Biometrika. 70 (1): 41–55. doi:10.2307/2335942.
- ↑ Lan, K. K. Gordon; DeMets, David L. (1983). "Discrete Sequential Boundaries for Clinical Trials". Biometrika. 70 (3): 659–663. doi:10.2307/2336502. JSTOR 2336502.
- ↑ Begaud B (1984). "Standardized assessment of adverse drug reactions: the method used in France. Special workshop—clinical". Drug Information Journal. 18 (3–4): 275–281. doi:10.1177/009286158401800314.
- ↑ "Revisiting Hurlbert 1984". Reflections on Papers Past. 29 November 2020. Retrieved 29 March 2022.
- ↑ Nuttin, Jozef M. Jr. (1985). "Narcissism beyond Gestalt and awareness: the name letter effect". European Journal of Social Psychology. 15 (3): 353–361. doi:10.1002/ejsp.2420150309.
- ↑ LaLonde, Robert (1986). "Evaluating the Econometric Evaluations of Training Programs with Experimental Data". American Economic Review. 4 (76): 604–620.
- ↑ Street, Anne Penfold; Street, Deborah J. (1987). "Combinatorics of Experimental Design". Clarendon Press.
- ↑ Hall, Marshall Jr. (January–February 1989), "Review of Combinatorics of Experimental Design", American Scientist, 77 (1): 91, JSTOR 27855619
- ↑ Wang, S. K.; Tsiatis, A. A. (1987). "Approximately optimal one-parameter boundaries for group sequential trials". Biometrics. 43 (1): 193–199. doi:10.2307/2531959. ISSN 0006-341X. JSTOR 2531959. PMID 3567304.
- ↑ Mead, R. (26 July 1990). The Design of Experiments: Statistical Principles for Practical Applications. Cambridge University Press. ISBN 978-0-521-28762-3.
- ↑ Latham, Gary P.; Erez, Miriam; Locke, Edwin A. (1988). "Resolving scientific disputes by the joint design of crucial experiments by the antagonists: Application to the Erez–Latham dispute regarding participation in goal setting". Journal of Applied Psychology. 73 (4): 753–772. doi:10.1037/0021-9010.73.4.753.
- ↑ He, Li (17 July 2003). "Design of Experiments Software, DOE software". The Chemical Information Network.
- ↑ Elkin, Irene; Shea, M. Tracie; Watkins, John T.; Imber, Stanley D.; Sotsky, Stuart M.; Collins, Joseph F.; Glass, David R.; Pilkonis, Paul A.; Leber, William R.; Docherty, John P.; Fiester, Susan J.; Parloff, Morris B. (November 1989). "National Institute of Mental Health Treatment of Depression Collaborative Research Program: General Effectiveness of Treatments". Archives of General Psychiatry. 46 (11): 971–982. doi:10.1001/archpsyc.1989.01810110013002. PMID 2684085.
{{cite journal}}:|access-date=requires|url=(help) - ↑ Haaland, Perry D. (1989). Experimental design in biotechnology. New York: Marcel Dekker. ISBN 9780824778811.
- ↑ Haaland, Perry D. (June 1991). "BOOK REVIEW: EXPERIMENTAL DESIGN IN BIOTECHNOLOGY Perry D. Haaland Marcel Dekkwe, Inc., New York, 1989". Drying Technology. 9 (3): 817–817. doi:10.1080/07373939108916715.
- ↑ Haaland, Perry D. (25 November 2020). "Experimental Design in Biotechnology". doi:10.1201/9781003065968.
{{cite journal}}: Cite journal requires|journal=(help) - ↑ Sacks, Jerome; Welch, William J.; Mitchell, Toby J.; Wynn, Henry P. (1989). "Design and Analysis of Computer Experiments". Statistical Science. 4 (4): 409–423. doi:10.1214/ss/1177012413.
- ↑ "Read "Statistical Methods for Testing and Evaluating Defense Systems: Interim Report" at NAP.edu". Retrieved 14 March 2021.
- ↑ Lam, C. W. H. (1991), "The Search for a Finite Projective Plane of Order 10", American Mathematical Monthly, 98 (4): 305–318, doi:10.2307/2323798, JSTOR 2323798
- ↑ Browne, Malcolm W. (20 December 1988), "Is a Math Proof a Proof If No One Can Check It?", The New York Times
- ↑ Tellner, Pär (29 June 2020). "ICH - Instilling 'harmony' for 30 years, and never more relevant than today!". European Federation of Pharmaceutical Industries and Associations (EFPIA). Retrieved 21 August 2026.
- ↑ Pearl, J. (1993). "Aspects of Graphical Models Connected With Causality". Proceedings of the 49th Session of the International Statistical Science Institute. pp. 391–401.
- ↑ "International Ethical Guidelines for Health-related Research Involving Humans" (PDF). Council for International Organizations of Medical Sciences (CIOMS), in collaboration with the World Health Organization. 2016. Retrieved 21 August 2026.
- ↑ Breggin, Ginger Ross; Breggin, Peter Roger (1995). Talking back to Prozac: what doctors won't tell you about today's most controversial drug. New York: St. Martin's Paperbacks. ISBN 978-0-312-95606-6.
- ↑ Card, David; Krueger, Alan B. (1994). "Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania". American Economic Review. 84 (4): 772–793. JSTOR 2118030.
- ↑ Neyer, Barry T. (February 1994). "A D-Optimality-Based Sensitivity Test". Technometrics. 36 (1): 61. doi:10.2307/1269199.
- ↑ Meyer, Ruth K.; Nachtsheim, Christopher J. (February 1995). "The Coordinate-Exchange Algorithm for Constructing Exact Optimal Experimental Designs". Technometrics. pp. 60–69.
- ↑ Kish, Leslie (1995). "Methods for design effects". Journal of Official Statistics. 11 (1): 55.
- ↑ Chaloner, Kathryn; Verdinelli, Isabella (1995), "Bayesian experimental design: a review", Statistical Science, 10 (3): 273–304, doi:10.1214/ss/1177009939
- ↑ "Integrated Addendum to ICH E6(R1): Guideline for Good Clinical Practice E6(R2)" (PDF). International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. 9 November 2016. Retrieved 31 March 2026.
- ↑ Jadad, A.R.; Moore R.A.; Carroll D.; Jenkinson C.; Reynolds D.J.M.; Gavaghan D.J.; McQuay H.J. (1996). "Assessing the quality of reports of randomized clinical trials: Is blinding necessary?". Controlled Clinical Trials. 17 (1): 1–12. doi:10.1016/0197-2456(95)00134-4. PMID 8721797.
- ↑ He, Li (17 July 2003). "Design of Experiments Software, DOE software". The Chemical Information Network.
- ↑ Begg, C.; Cho, M.; Eastwood, S.; Horton, R.; Moher, D.; Olkin, I.; Pitkin, R.; Rennie, D.; Schulz, K.F.; Simel, D.; Stroup, D.F. (1996). "Improving the quality of reporting of randomized controlled trials: the CONSORT statement". JAMA. 276 (8): 637–639.
- ↑ Kienle, Gunver S.; Kiene, Helmut (December 1997). "The powerful placebo effect: fact or fiction?". Journal of Clinical Epidemiology. 50 (12): 1311–1318. doi:10.1016/s0895-4356(97)00203-5. PMID 9449934.
{{cite journal}}:|access-date=requires|url=(help) - ↑ Mitchell, T. (1997). Machine Learning. New York: McGraw Hill. ISBN 0-07-042807-7.
- ↑ Shalala, D.E. (26 August 1998). "Policy statement on changing the population standard used for age adjusting death rates in DHHS publications". Retrieved 13 August 2026.
- ↑ UK Prospective Diabetes Study Group (1998). "Tight Blood Pressure Control and Risk of Macrovascular and Microvascular Complications in Type 2 Diabetes: UKPDS 38". BMJ: British Medical Journal. 317 (7160): 703–713. doi:10.1136/bmj.317.7160.703. JSTOR 25180360.
- ↑ Horne, Gary; Schwierz, Klaus-Peter (1 March 2016). "Summary of Data Farming". Axioms. MDPI. p. 8. Retrieved 21 August 2026.
- ↑ Dehejia, R.H.; Wahba, S. (1999). "Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs". Journal of the American Statistical Association. 94 (448): 1053–1062.
- ↑ Kramarz, Piotr, Eric K. France, Frank Destefano, Steven B. Black, Henry Shinefield, Joel I. Ward, Emily J. Chang et al. "Population-based study of rotavirus vaccination and intussusception." The Pediatric Infectious Disease Journal 20, no. 4 (2001): 410-416.
- ↑ "Analyzing Families of Experiments in SE: a Systematic Mapping Study" (PDF). arxiv.org. Retrieved 12 March 2022.
- ↑ Oehlert, Gary W. (2010). A First Course in Design and Analysis of Experiments. Gary W. Oehlert.
- ↑ Ahmad, O.B.; Boschi-Pinto, C.; Lopez, A.D.; Murray, C.J.L.; Lozano, R.; Inoue, M. (2001). "Age Standardization of Rates: A New WHO Standard" (PDF). GPE Discussion Paper Series: No.31. World Health Organization.
- ↑ "Adversarial Collaboration: An EDGE Lecture by Daniel Kahneman | Edge.org". www.edge.org. Retrieved 8 March 2022.
- ↑ Berger, Michele W.; (University of Pennsylvania). "In the pursuit of scientific truth, working with adversaries can pay off". phys.org. Retrieved 13 August 2026.
- ↑ Devereaux, P.J.; Manns, Braden J.; Ghali, William A.; Quan, Hude; Lacchetti, Christina; Montori, Victor M.; Bhandari, Mohit; Guyatt, Gordon H. (April 2001). "Physician interpretations and textbook definitions of blinding terminology in randomized controlled trials". JAMA. 285 (15): 2000–2003. doi:10.1001/jama.285.15.2000. PMID 11308438.
{{cite journal}}:|access-date=requires|url=(help) - ↑ Moher, D.; Schulz, K.F.; Altman, D.G. (2001). "The CONSORT statement: revised recommendations for improving the quality of reports of parallel-group randomized trials". Annals of Internal Medicine. 134 (8): 657–662.
- ↑ "WMA - Policy". Archived from the original on 20 February 2009.
- ↑ Musch, J., & Klauer, K. C. (2002). Psychological experimenting on the World Wide Web: Investigating content effects in syllogistic reasoning. In B. Batinic, U.-D. Reips, & M. Bosnjak (Eds.), Online social sciences (pp. 181–212). Hogrefe & Huber Publishers.
- ↑ Reips, U.-D. (2002). Standards for internet-based experimenting. Experimental Psychology, 49(4), 243–256.
- ↑ Bloom, H.S.; Michalopoulos, C.; Hill, C.J.; Lei, Y. (2002). "Can Nonexperimental Comparison Group Methods Match the Findings from a Random Assignment Evaluation of Mandatory Welfare-to-Work Programs?". MDRC Working Papers on Research Methodology.
{{cite web}}: Missing or empty|url=(help) - ↑ Schneider, ed. by Sandra L.; Shanteau, James (2003). Emerging perspectives on judgment and decision research. Cambridge [u.a.]: Cambridge Univ. Press. pp. 438–9. ISBN 052152718X.
{{cite book}}:|first=has generic name (help) - ↑ "Randomized controlled trials in development economics". Wikipedia. Retrieved 3 April 2026.
- ↑ Camerer, Colin (2003). Behavioral Game Theory: Experiments in Strategic Interaction. Princeton University Press. p. 42. ISBN 978-0691090399.
- ↑ Hirata, S. (2003). "Cooperation in chimpanzees". Hattatsu. 95: 103–111.
- ↑ Abadie, Alberto; Gardeazabal, Javier (2003). "The Economic Costs of Conflict: A Case Study of the Basque Country". American Economic Review. 93 (1): 113–132. doi:10.1257/000282803321455188.
- ↑ "Critical Path Innovation Meetings: Guidance for Industry". U.S. Food and Drug Administration, Center for Drug Evaluation and Research. April 2015.
- ↑ Campbell MK, Elbourne DR, Altman DG (2004). "CONSORT statement: extension to cluster randomised trials". BMJ. 328 (7441): 702–708. doi:10.1136/bmj.328.7441.702. PMC 381234. PMID 15031246.
- ↑ Pildal J, Chan AW, Hróbjartsson A, Forfang E, Altman DG, Gøtzsche PC (2005). "Comparison of descriptions of allocation concealment in trial protocols and the published reports: cohort study". BMJ. 330 (7499): 1049. doi:10.1136/bmj.38414.422650.8F. PMC 557221. PMID 15817527.
- ↑ Pocock S (2005). "When (not) to stop a clinical trial for benefit" (PDF). JAMA. 294 (17): 2228–2230. doi:10.1001/jama.294.17.2228. PMID 16264167.
- ↑ Ioannidis, John P. A. (2005-08-30). "Why Most Published Research Findings Are False". PLOS Medicine. 2 (8): e124. doi:10.1371/journal.pmed.0020124. PMC 1182327. PMID 16060722.
- ↑ Kleijnen, J.P.C.; Sanchez, S.M.; Lucas, T.W.; Cioppa, T.M. (2005). "A User's Guide to the Brave New World of Designing Simulation Experiments". INFORMS Journal on Computing. 17 (3): 263–289. doi:10.1287/ijoc.1050.0136.
- ↑ Swan M (June 2013). "The Quantified Self: Fundamental Disruption in Big Data Science and Biological Discovery". Big Data. 1 (2): 85–99. doi:10.1089/big.2012.0002. PMID 27442063.
- ↑ "Exploratory IND Studies, Guidance for Industry, Investigators, and Reviewers" (PDF). Food and Drug Administration. January 2006.
- ↑ Haahr, M.T.; Hróbjartsson, A. (2006). "Who is blinded in randomized clinical trials? A study of 200 trials and a survey of authors". Clinical Trials. 3 (4): 360–365. doi:10.1177/1740774506069153. PMID 17060219.
- ↑ Melis, Alicia P.; Hare, Brian; Tomasello, Michael (2006). "Engineering cooperation in chimpanzees: Tolerance constraints on cooperation". Animal Behaviour. 72 (2): 275–286. doi:10.1016/j.anbehav.2005.09.018.
- ↑ Wisely, Janet (2007). "Building on Improvement: Establishing a National Research Ethics Service". Research Ethics. 3: 3–4. doi:10.1177/174701610700300102. S2CID 167972296.
- ↑ McCrary (2008). "Manipulation of the Running Variable in the Regression Discontinuity Design: A Density Test". Journal of Econometrics. 142 (2): 698–714. doi:10.1016/j.jeconom.2007.05.005.
- ↑ Jones, Bradley; Lin, Dennis K. J.; Nachtsheim, Christopher J. (January 2008). "Bayesian D-optimal supersaturated designs". Journal of Statistical Planning and Inference. pp. 86–92.
- ↑ Wood, L; Egger, M; Gluud, LL; Schulz, KF; Jüni, P; Altman, DG; Gluud, C; Martin, RM; Wood, AJ; Sterne, JA (2008). "Empirical evidence of bias in treatment effect estimates in controlled trials with different interventions and outcomes: meta-epidemiological study". BMJ. 336 (7644): 601–605. doi:10.1136/bmj.39465.451748.AD. PMC 2267990. PMID 18316340.
- ↑ Kahneman, Daniel; Klein, Gary. Conditions for intuitive expertise: A failure to disagree. American Psychologist, Vol 64(6), Sep 2009, 515-526. doi: 10.1037/a0016755
- ↑ Wagenmakers, E.-J., Wetzels, R., Borsboom, D., & van der Maas, H. L. J. (2010). Why psychologists must change the way they analyze their data: The case of psi.
- ↑ Hróbjartsson A, Gøtzsche PC (January 2010). Hróbjartsson A (ed.). "Placebo interventions for all clinical conditions" (PDF). The Cochrane Database of Systematic Reviews. 106 (1): CD003974. doi:10.1002/14651858.CD003974.pub3. PMID 20091554.
- ↑ Kilkenny, Carol; Browne, William J.; Cuthill, Innes C.; Emerson, Michael; Altman, Douglas G. (29 June 2010). "Improving Bioscience Research Reporting: The ARRIVE Guidelines for Reporting Animal Research". PLOS Biology. 8 (6) e1000412. doi:10.1371/journal.pbio.1000412. PMC 2893951. PMID 20613859.
{{cite journal}}: CS1 maint: unflagged free DOI (link) - ↑ Kilkenny, Carol; Parsons, Nick; Kadyszewski, Ed; Festing, Michael F. W.; Cuthill, Innes C.; Fry, Derek; Hutton, Jane; Altman, Douglas G. (30 November 2009). "Survey of the Quality of Experimental Design, Statistical Analysis and Reporting of Research Using Animals". PLOS ONE. 4 (11) e7824. doi:10.1371/journal.pone.0007824. PMC 2779358. PMID 19956596.
{{cite journal}}: CS1 maint: unflagged free DOI (link) - ↑ Center for Drug Evaluation and Research (CDER); Center for Biologics Evaluation and Research (CBER) (February 2010). "Adaptive Design Clinical Trials for Drugs and Biologics" (PDF). U.S. Food and Drug Administration.
- ↑ Malone, J.; Holloway, E.; Adamusiak, T.; Kapushesky, M.; Zheng, J.; Kolesnikov, N.; Zhukova, A.; Brazma, A.; Parkinson, H. (2010). "Modeling sample variables with an Experimental Factor Ontology". Bioinformatics. 26 (8): 1112–1118. doi:10.1093/bioinformatics/btq099. PMC 2853691. PMID 20200009.
- ↑ Abadie, Alberto; Diamond, Alexis; Hainmueller, Jens (2010). "Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California's Tobacco Control Program". Journal of the American Statistical Association. 105 (490): 493–505. doi:10.1198/jasa.2009.ap08746. hdl:1721.1/59447.
- ↑ Schulz, K.F.; Altman, D.G.; Moher, D. (2010). "CONSORT 2010 statement: updated guidelines for reporting parallel group randomised trials". BMJ. 340 c332.
- ↑ Wang, Shirley S. (30 December 2013). "Health: Scientists Look to Improve Cost and Time of Drug Trials". Wall Street Journal. Retrieved 13 August 2026.
- ↑ Jones, Bradley; Nachtsheim, Christopher J. (2011). "A Class of Three-Level Designs for Definitive Screening in the Presence of Second-Order Effects". Journal of Quality Technology. pp. 1–15.
- ↑ Iacus, Stefano M.; King, Gary; Porro, Giuseppe (2011). "Multivariate Matching Methods That Are Monotonic Imbalance Bounding". Journal of the American Statistical Association. 106 (493): 345–361. doi:10.1198/jasa.2011.tm09599. hdl:2434/151476.
- ↑ Deng, Alex; Xu, Ya; Kohavi, Ron; Walker, Toby (2013). "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data". WSDM 2013: Sixth ACM International Conference on Web Search and Data Mining.
- ↑ Simonsohn, Uri; Nelson, Leif D.; Simmons, Joseph P. (2014). "P-curve: A key to the file-drawer". Journal of Experimental Psychology: General. 143 (2): 534–547. doi:10.1037/a0033242. PMID 23855496.
- ↑ Keevash, Peter (2014). "The existence of designs". arXiv preprint.
- ↑ Gelman, A; Carlin, J (2014). "Beyond power calculations: Assessing Type S (sign) and Type M (magnitude) errors". Perspectives in Psychological Science. 9 (6): 641–651. doi:10.1177/1745691614551642.
- ↑ Walsh, Michael; Srinathan, Sadeesh K.; McAuley, Daniel F.; Mrkobrada, Marko; Levine, Oren; Ribic, Christine; Molnar, Amit O.; Dattani, Nishith D.; Burke, Andrew; Guyatt, Gordon; Thabane, Lehana; Walter, Stephen D.; Pogue, Janice; Devereaux, P.J. (June 2014). "The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index". Journal of Clinical Epidemiology. 67 (6): 622–628. doi:10.1016/j.jclinepi.2013.10.019. PMID 24508144. Retrieved 29 August 2026.
- ↑ Bohannon, John (27 May 2015). "I Fooled Millions Into Thinking Chocolate Helps Weight Loss. Here's How". Gizmodo. Retrieved 13 August 2026.
- ↑ Head, Megan L.; Holman, Luke; Lanfear, Rob; Kahn, Andrew T.; Jennions, Michael D. (2015-03-13). "The Extent and Consequences of P-Hacking in Science". PLOS Biology. 13 (3) e1002106. doi:10.1371/journal.pbio.1002106. PMC 4359000. PMID 25768323.
- ↑ Miller, Claire Cain (25 February 2016). "Is Blind Hiring the Best Hiring?". The New York Times. Retrieved 13 August 2026.
- ↑ Athey, Susan, and Guido Imbens (2016), "Recursive partitioning for heterogeneous causal effects." Proceedings of the National Academy of Sciences 113, (27), 7353–7360.
- ↑ "Promoting reproducibility with registered reports". Nature Human Behaviour. 10 January 2017. Retrieved 21 August 2026.
- ↑ "International Collaborative Network for N-of-1 Trials and Single-Case Designs". N-of-1 and SCED. Retrieved 2024-07-09.
- ↑ Aronow, Peter M.; Samii, Cyrus (2017-12-01). "Estimating average causal effects under general interference, with application to a social network experiment". The Annals of Applied Statistics. 11 (4): 1912–1947. arXiv:1305.6156. doi:10.1214/16-aoas1005.
- ↑ Kennedy, Andrew D. M.; Torgerson, David J.; Campbell, Marion K.; Grant, Adrian M. (2017). "Subversion of allocation concealment in a randomised controlled trial: a historical case study". Trials. 18 (1): 204. doi:10.1186/s13063-017-1946-z. PMC 5414185. PMID 28464922.
{{cite journal}}: CS1 maint: unflagged free DOI (link) - ↑ Saltaji, Humam; Armijo-Olivo, Susan; Cummings, Greta G.; Amin, Maryam; da Costa, Bruno R.; Flores-Mir, Carlos (18 May 2018). "Influence of blinding on treatment effect size estimate in randomized controlled trials of oral health interventions". BMC Medical Research Methodology. 18: 42. doi:10.1186/s12874-018-0491-0. PMC 5960173. PMID 29776394. Retrieved 29 August 2026.
{{cite journal}}: CS1 maint: unflagged free DOI (link) - ↑ Gordon, Brett R.; Zettelmeyer, Florian; Bhargava, Neha; Chapsky, Dan (2018). "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook".
- ↑ Nosek, Brian A.; Ebersole, Charles R.; DeHaven, Alexander C.; Mellor, David T. (2018-03-13). "The preregistration revolution". Proceedings of the National Academy of Sciences. 115 (11): 2600–2606. doi:10.1073/pnas.1708274114.
- ↑ Yeh, Robert W.; Valsdottir, Linda S.; Yeh, Molly W.; Shen, Changyu; Kramer, Daniel B.; Strom, Jordan B.; Secemsky, Eric A.; Healy, Jordan L.; Domeier, Robert M.; Kazi, Dhruv S.; Nallamothu, Brahmajee K. (13 December 2018). "Parachute use to prevent death and major trauma when jumping from aircraft: randomized controlled trial". BMJ. 363: k5094. doi:10.1136/bmj.k5094. PMID 30545967.
{{cite journal}}:|access-date=requires|url=(help) - ↑ Schulz, KF; Chalmers, I; Altman, DG; Grimes, DA; Moher, D; Hayes, RJ (June 2018). "'Allocation concealment': the evolution and adoption of a methodological term". Journal of the Royal Society of Medicine. 111 (6): 216–224. doi:10.1177/0141076818776604. PMC 6022887. PMID 29877772.
- ↑ "Nobel Prize 2019 experimental approach to poverty". Wikipedia. Retrieved 3 April 2026.
- ↑ "Adaptive designs for clinical trials of drugs and biologics: Guidance for industry". U.S. Food and Drug Administration (FDA). 1 November 2019. Retrieved 7 April 2021.
- ↑ King, Gary; Nielsen, Richard (October 2019). "Why Propensity Scores Should Not Be Used for Matching". Political Analysis. 27 (4): 435–454. doi:10.1017/pan.2019.11. hdl:1721.1/128459.
- ↑ Fabijan, Aleksander; Gupchup, Jayant; Gupta, Somit; Omhover, Jeff; Qin, Wen; Vermeer, Lukas; Dmitriev, Pavel (2019-07-25). "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments". Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. pp. 2156–2164. doi:10.1145/3292500.3330722. ISBN 978-1-4503-6201-6.
- ↑ Jones, Bradley; Lekivetz, Ryan; Majumdar, Dibyen; Nachtsheim, Christopher J.; Stallrich, Jonathan W. (August 2020). "Construction, Properties, and Analysis of Group-Orthogonal Supersaturated Designs". Technometrics. pp. 403–414.
- ↑ Kotok, Alan (19 March 2020). "WHO beginning Covid-19 therapy trial". Technology News: Science and Enterprise. Retrieved 7 April 2020.
- ↑ "Launch of a European clinical trial against COVID-19". INSERM. 22 March 2020. Retrieved 5 April 2020.
- ↑ Sert, Nathalie Percie du; Hurst, Viki; Ahluwalia, Amrita; Alam, Sabina; Avey, Marc T.; Baker, Monya; Browne, William J.; Clark, Alejandra; Cuthill, Innes C.; Dirnagl, Ulrich; Emerson, Michael (14 July 2020). "The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research". PLOS Biology. 18 (7) e3000410. doi:10.1371/journal.pbio.3000410. PMC 7360023. PMID 32663219.
{{cite journal}}: CS1 maint: unflagged free DOI (link) - ↑ O'Grady, Cathleen (14 July 2020). "Journals endorse new checklist to clean up sloppy animal research". Science. Retrieved 13 August 2026.
- ↑ Garcia-Pelegrin, Elias; Schnell, Alexandra K.; Wilkins, Clive; Clayton, Nicola S. (2020). "An unexpected audience". Science. 369 (6510): 1424–1426. doi:10.1126/science.abc6805. PMID 32943508.
- ↑ Adda, Jérôme; Decker, Christian; Ottaviani, Marco (2020-06-16). "P-hacking in clinical trials and how incentives shape the distribution of results across phases". Proceedings of the National Academy of Sciences of the United States of America. 117 (24): 13386–13392. arXiv:1907.00185. doi:10.1073/pnas.1919906117. PMID 32487730.
- ↑ Busby, Mattha (27 June 2020). "Cambridge college to remove window commemorating eugenicist". The Guardian. Retrieved 2020-06-28.
- ↑ Erslev, Malthe Stavning (2022). "A Mimetic Method: Rendering Artificial Intelligence Imaginaries through Enactment". A Peer-Reviewed Journal About. 11 (1): 34–49. doi:10.7146/aprja.v11i1.134305.
- ↑ Garcia-Pelegrin, Elias; Schnell, Alexandra K.; Wilkins, Clive; Clayton, Nicola S. (2021). "Exploring the perceptual inabilities of Eurasian jays (Garrulus glandarius) using magic effects". PNAS. 118 (24) e2026106118. doi:10.1073/pnas.2026106118. PMC 8214664. PMID 34074798.
- ↑ Clark, Cory J.; Costello, Thomas; Mitchell, Gregory; Tetlock, Philip E. (March 2022). "Keep your enemies close: Adversarial collaborations will improve behavioral science". Journal of Applied Research in Memory and Cognition. 11 (1): 1–18. doi:10.1037/mac0000004.
- ↑ "In the pursuit of scientific truth, working with adversaries can pay off". Penn Today. 7 July 2022. Retrieved 13 August 2026.
- ↑ Toews, Ingrid; Anglemyer, Andrew; Nyirenda, John Lz; Alsaid, Dima; Balduzzi, Sara; Grummich, Kathrin; Schwingshackl, Lukas; Bero, Lisa (2024-01-04). "Healthcare outcomes assessed with observational study designs compared with those assessed in randomized trials: a meta-epidemiological study". The Cochrane Database of Systematic Reviews. 1 (1): MR000034. doi:10.1002/14651858.MR000034.pub3. PMID 38174786.
- ↑ Feng, Tianxing; Zhu, Chouwen; Yang, Geliang (2025). "Viewpoints of investigator on CONSORT 2025 statement-updated guideline for reporting randomized trials". Journal of Thoracic Disease. 17 (5): 2752–2755. doi:10.21037/jtd-2025-871.
{{cite journal}}: CS1 maint: unflagged free DOI (link) - ↑ Msaouel, Pavlos (2025-03-28). "The curious rise of randomised non-comparative trials". Significance. 22 (3): 40–44. doi:10.1093/jrssig/qmaf029. ISSN 1740-9705.
- ↑ Sherry, Alexander D.; Msaouel, Pavlos; Ludmir, Ethan B. (2024). "A meta-epidemiological analysis of post-hoc comparisons and primary endpoint interpretability among randomized noncomparative trials in clinical medicine". Journal of Clinical Epidemiology. 175 111540. doi:10.1016/j.jclinepi.2024.111540.
- ↑ Cashin, Aidan G.; Hansford, Harrison J.; Hernán, Miguel A.; Swanson, Sonja A.; Lee, Hopin; Jones, Matthew D.; Dahabreh, Issa J.; Dickerman, Barbra A.; Egger, Matthias; Garcia-Albeniz, Xabier; Golub, Robert M.; Islam, Nazrul; Lodi, Sara; Moreno-Betancur, Margarita; Pearson, Sallie-Anne; Schneeweiss, Sebastian; Sharp, Melissa K.; Sterne, Jonathan A. C.; Stuart, Elizabeth A.; McAuley, James H. (2025-09-03). "Transparent Reporting of Observational Studies Emulating a Target Trial-The TARGET Statement". JAMA. 334 (12): 1084–1093. doi:10.1001/jama.2025.13350.
- ↑ "Image-based treatment effect heterogeneity". arXiv. Retrieved 3 April 2026.
- ↑ "Design of experiments". Google Trends. Retrieved 14 December 2021.
- ↑ "Design of experiments". books.google.com. Retrieved 14 December 2021.
- ↑ "Design of experiments". wikipediaviews.org. Retrieved 14 December 2021.