Timeline of experiment design

From Timelines
Jump to navigation Jump to search

This is a timeline of experiment design, attempting to describe significant and illustrative events in the history of the field.

Sample questions

The following are some interesting questions that can be answered by reading this timeline:

Big picture

Time period Development summary More details
Ancient Times – 1700s Early foundations Experimental thinking has ancient roots, notably in medicine and agriculture. One of the earliest recorded episodes resembling a clinical trial dates to around 500 BCE, when King Nebuchadnezzar compares a meat-and-wine diet against one of beans and water, observing the latter group appears healthier.[1] Greek physicians like Hippocrates[2] and Galen emphasize observation and rudimentary comparative methods. However, formal experimentation remains rare. The Scientific Revolution in the 1600s shifts focus from authority to empirical investigation. Francis Bacon promotes inductive reasoning and systematic observation, laying philosophical groundwork for modern science.[3] Nonetheless, experiments lack standardized design or statistical rigor. Scientists like Galileo[4] and Boyle[5] use controlled observations, but their methods are more qualitative than quantitative. The period ends with the emergence of probability theory by Pascal[6], Fermat[7], and Bernoulli[8], which begins to offer tools later essential for experimental statistics.
1700s – Late 1800s Classical period The Age of Enlightenment brings advances in both scientific methodology and mathematical statistics.[9] Controlled experiments become more common in chemistry, physics, and agriculture. Notably, James Lind conducts one of the first controlled clinical trials in 1747, testing treatments for scurvy.[10] In parallel, Carl Linnaeus systematizes biological classification, influencing observational rigor.[11] Pierre-Simon Laplace and Carl Friedrich Gauss develop early statistical tools such as the normal distribution and least squares estimation, crucial for data analysis. Agricultural trials, especially in Europe, begin to use basic comparison and replication. However, randomness, replication, and formal control are not yet standardized. This period sets the stage for modern experimentation by linking empirical testing with early statistical theory, but experiments still lack formal design structures and reproducibility standards. The idea of a placebo effect—a therapeutic outcome derived from an inert treatment— is already discussed in 18th century psychology[12]
Early 1900s – 1950s Formalization and statistics This period marks the birth of modern experimental design. Ronald A. Fisher revolutionizes the field in the 1920s and 1930s with work at Rothamsted Experimental Station in England.[13] Fisher introduces core principles: randomization, replication, and blocking. He develops the analysis of variance (ANOVA) and emphasizes the role of statistical inference in drawing conclusions. His 1935 book, The Design of Experiments, would remain foundational. Experimental design becomes central to agriculture[14][14][15], biology, and later psychology and social sciences. Jerzy Neyman and Egon Pearson introduce hypothesis testing frameworks, further refining experimental methodology.[16] This era institutionalizes statistics as an essential scientific tool and formalizes the structure of experiments, enabling reproducibility and greater objectivity in scientific inference. The need to blind researchers becomes widely recognized in the mid-century.[17]
1960s – present Modern and computational era Post-1950s, experimental design expand into diverse fields: psychology, economics, medicine, and engineering. With the rise of computers, simulation, factorial designs, and response surface methodology become widespread. Clinical trials adopt strict protocols, often randomized and double-blind, guided by ethical oversight. In social sciences, randomized controlled trials (RCTs) gain ground, supported by behavioral economics and development studies. The Bayesian approach gains traction, offering alternative inference frameworks. From the 1990s, machine learning and A/B testing transform experimental design in tech industries, allowing real-time, large-scale experimentation. Today, experimental design is deeply intertwined with data science, AI, and evidence-based policy, continually evolving to address high-dimensional data, ethical concerns, and reproducibility challenges across disciplines.
Time period Development summary More details
1880s Method A solution to the scheduling of round-robin tournaments — arranging matches so that each competitor plays every other competitor exactly once across a series of rounds — is published, connecting practical tournament-scheduling problems to the combinatorial theory underlying balanced designs.[18] Global
1910s Concept Ronald A. Fisher develops the concept of variance as a measure of population variability, laying the foundation for later statistical methods such as analysis of variance (ANOVA).[19] United Kingdom
1920s Method Ronald Fisher develops randomized block design and formal principles of blocking to control variation in experiments. United Kingdom
1920s Method Fisher introduces key principles of modern experimental design, including confounding, randomization, replication, blocking, Latin square designs, and other factorial designs.[19] United Kingdom
1920s Method analysis of variance and maximum likelihood estimation are developed as major statistical tools for analyzing experimental data.[19] United Kingdom
1920s Organization The first academic statistics department is established at Iowa State University under George Snedecor, helping institutionalize statistical experimental methods.[19] United States
1920s Method George Snedecor develops the F-test and F-distribution, which become central tools for decisions in ANOVA-based experimental designs.[19] United States
1930s Concept The size of an experiment is recognized as a key design decision, with sample size linked to precision, significance, and statistical power.[20] Global
1930s Concept The advantages of factorial experimentation over one-factor-at-a-time methods are articulated, especially its economy of resources and ability to estimate interrelationships among factors.[20] Global
1930s Field development Early applications of statistical experimental design begin to appear in industrial settings, extending beyond agriculture.[21] Global
1930s Method Randomization is formalized as the use of random numbers or equivalent devices to assign treatments and reduce bias from extraneous variation.[20] Global
1930s Concept Experimental error is formally distinguished as a major obstacle to precision, arising from extraneous sources of variation beyond the treatments themselves.[20] Global
1930s Concept Bias is distinguished from random error, emphasizing that systematic error can mislead conclusions and often cannot be detected through statistical analysis alone.[20] Global
1930s Method double-blind experiments and placebo controls are recognized as important safeguards against bias in experiments involving subjective or clinical judgments.[20] Global
1930s Method Refinements of experimental technique, such as practice runs, clearer instructions, environmental control, and improved measurement instruments, are emphasized as ways to reduce experimental error.[20] Global
1930s Method blocking is presented as a method for increasing experimental precision by balancing treatments across subjects or units with similar initial characteristics.[20] Global
1930s Method matched pairs design and matched groups designs are used with human subjects as forms of blocking analogous to blocks in agricultural experiments.[20] Global
1930s Concept Core principles of experimental design—randomization, replication, and blocking—are established as fundamental techniques to reduce bias and improve precision in experiments.[22] United Kingdom
1930s Method analysis of covariance is introduced as a more accurate statistical adjustment for initial differences when blocking is not feasible or when further precision is needed.[20] Global
1930s Method factorial experimentation is formalized as a way to investigate several factors simultaneously, improving efficiency and enabling study of interactions among variables.[20] Global
1940s Method George Box and collaborators develop early central composite design methods for response surface optimization. United Kingdom
1940s Method sequential analysis emerges during World War II to improve efficiency in military experimentation and decision-making.[14] United States
1940s Method Development of orthogonal designs and Latin squares as structured experimental layouts for controlling variation in experiments.[14] United Kingdom
Early 1950s Concept C. R. Rao introduces the concept of orthogonal arrays as a formal framework for experimental designs, building on and generalizing the efficient main-effect designs found by Raj Chandra Bose and K. Kishen in 1940. Around this time, Genichi Taguchi visits the Indian Statistical Institute, where he encounters this concept; it goes on to play a central role in his later development of the Taguchi methods, later adopted by Japanese and Indian industries and eventually, with some reservations, by US industry.[23] India
1950s Method W. Edwards Deming promotes statistical quality control in Japan, initiating the quality revolution and widespread industrial use of experimental design.[14] Japan
1950s Method Statistical methods and experimental design spread rapidly through the medical literature, extending techniques developed for agriculture into biomedical research.[19] Global
1950s Method Factorial experimental designs become widely used in industrial experimentation to study multiple variables simultaneously United States
1950s Method Control groups are emphasized as essential in evaluatory studies for distinguishing treatment effects from other changes over time.[20] Global
1950s–1970s Field development Statistical experimental design expands from agriculture into industrial applications, particularly in chemical and process industries.[24] Global
1950s–1980s Field development Response surface methods and statistical design techniques spread across chemical and process industries, especially in research and development.[21] Global
Mid-20th century Method Standardized procedural steps for experimental design are established, including problem definition, variable selection, factor identification, experimental execution, and statistical analysis.[22] Global}
Mid-20th century Method analysis of variance (ANOVA) becomes the central statistical framework for analyzing experimental data and guiding experimental design decisions.[22] Global
1960s Method Japanese industries adopt experimental design and statistical quality control, leading to major improvements in manufacturing quality and global competitiveness.[14] Japan
1960s Experiment randomized controlled trial becomes the gold standard in clinical research, replacing anecdotal medical evidence with controlled experimental methods.[14] United States
1960s–1970s Concept Industrial experiments are recognized as distinct due to immediacy and sequential learning, enabling rapid iteration and adaptive experimentation.[21] Global
1970s Method Multi center clinical trials become standard, enabling large-scale randomized evaluation across diverse populations Global [25]
1970s–1980s Method Genichi Taguchi develops robust parameter design methods to reduce variability and improve product quality.[14] Japan
1970s–1990 Method Quality improvement methodologies such as Total Quality Management (TQM) and Continuous Quality Improvement (CQI) integrate experimental design into industrial processes.[14] Global
Late 1970s Method Genichi Taguchi promotes robust parameter design, emphasizing quality improvement and reduction of variability using orthogonal arrays and fractional factorial designs.[24] Japan
Late 1970s Method Genichi Taguchi introduces robust parameter design, emphasizing reduction of variability and improvement of product quality under noise conditions.[21] Japan
Late 1970s–1980s Field development Growing interest in quality improvement in Western industry leads to broader adoption of experimental design methods.[21] Global
1980s Method Quality engineering approaches emphasize robustness and variance reduction in manufacturing processes Japan [26]
1980s Method Meta analysis emerges as a method to systematically combine results from multiple experiments Global
1980s Field development Designed experiments become widely adopted in manufacturing industries (automotive, aerospace, electronics), driven by quality improvement initiatives.[24] Global
1980s Method Large-scale randomized trials such as the ISIS heart studies demonstrate the feasibility of very large clinical experiments.[27] Global
Late 1980s Concept Critical evaluation of Taguchi methods leads to refinement of statistical foundations and development of alternative robust design methodologies.[24] Global
1990s Method Six Sigma integrates statistical experimental design into industrial quality control and process improvement,[14] United States
1990s Method optimal design approaches gain prominence, focusing on maximizing information and efficiency under resource constraints, supported by advances in computing.[24] Global
1990s Field development Designed experiments expand into new domains including services, finance, and government operations.[24] Global
1990s Method The modern era of experimental design begins, driven by globalization and economic competition, expanding DOE across industries and services.[14] Global
1990s–present Field development The Optimal Design Era emerges, characterized by widespread use of computational algorithms, software tools, and expansion of DOE across diverse sectors.[24] Global
2000s Method Online randomized controlled experiments (A/B testing) become standard in internet companies for product and interface optimization, enabling continuous data-driven decision making.[28] United States
2000s Method Bradley Jones and Chris Nachtsheim develop definitive screening designs, enabling efficient identification of important factors with minimal runs.[24] United States
2000s Method Bradley Jones introduces Bayesian D-optimal design approaches, incorporating prior knowledge into experimental planning.[24] United States
2000s Method Bradley Jones develops the Custom Design platform, enabling flexible, computer-generated optimal experimental designs tailored to specific problems.[24] United States
2000s Method Advances in supersaturated designs and split-plot designs improve efficiency in experiments with many factors and constraints.[24] United States
2000s–present Method Development of group-orthogonal supersaturated designs (GO-SSDs) improves factor screening when variables exceed experimental runs.[24] United States
2010s Method Continuous A/B testing systems enable rapid, large-scale experimentation in digital platforms, allowing ongoing optimization based on user behavior data.[29] Global
Late 20th century Application Experimental design expands beyond agriculture into engineering, business, and scientific research, becoming a general-purpose methodology.[22] Global

Full timeline

Year Event type Details Geographical location
ca. 500 BCE Experiment Around this time, what would be often described as the earliest recorded experiment resembling a clinical trial occurs when King Nebuchadnezzar compares two diets—meat and wine versus beans and water—and observes that the group consuming the vegetarian diet appears healthier.[1] Babylon (Mesopotamia)
ca. 500 BCE Experiment According to the Book of Daniel, Daniel and his companions request a ten-day trial in which they consume only vegetables and water, while other youths at the royal court continue eating the king's rich diet; the results of the two regimens are compared, and the steward acts based on the observable outcome. Several scholars of ancient literature and the history of science have cited this account as one of the earliest descriptions of a controlled comparative test, and biblical historian David C. Lindberg specifically includes it among ancient examples where a trial is deliberately arranged to evaluate competing conditions.[30][31][32] Babylon (Mesopotamia)
ca. 400 BCE Concept Hippocrates promotes systematic observation and naturalistic explanations in medicine, emphasizing empirical study over supernatural interpretations.[1] Greece
ca. 587 AD Concept Indian astronomer-mathematician Varahamihira, in his text Brhat Samhita, describes what is considered one of the earliest datable applications of combinatorial design: a method for making perfumes by selecting 4 substances from 16 available substances using a magic square arrangement.[33] India
c. 1021 Method Arab polymath Ibn al-Haytham completes the Book of Optics, applying a methodical, inductive-experimental approach to optical and mathematical problems inherited from Ptolemy — emphasizing self-criticality, reliance on visible experimental results over inherited authority, and rigorous skepticism toward earlier scholars' conclusions. He explicitly instructs that a researcher should "make himself an enemy of all that he reads" and interrogate it from every angle, while remaining alert to his own susceptibility to prejudice and leniency. Historians of science regard Ibn al-Haytham as one of the first scholars to systematically apply controlled, inductive experimentation to derive knowledge, prefiguring elements of the modern scientific method by roughly six centuries.[34] Egypt
1025 AD Concept Persian polymath Avicenna completes The Canon of Medicine, a major medical encyclopedia synthesizing Greco-Roman and Islamic knowledge. The work describes diseases, treatments, and hundreds of drugs, and outlines principles for experimentally testing medicines, including reproducibility of results, shaping medical education in both the Islamic world and Europe for centuries.[35] Persia
1620 Concept Francis Bacon publishes Novum Organum, proposing a new empirical approach to scientific inquiry. He argues that knowledge should be based on systematic observation and inductive reasoning rather than tradition or pure logic, helping establish methodological foundations for the modern scientific method and experimental science.[36] England
1676 Experiment Danish astronomer Ole Rømer demonstrates that light does not travel instantaneously but has a measurable finite speed, by observing that the apparent timing of Jupiter's moons' eclipses was delayed when Jupiter was farther from Earth compared to when it was closer — an early landmark example of a natural experiment, relying on observation of a system as it naturally occurs rather than direct manipulation of variables.[37] Denmark
1683 Method Robert Boyle solidifies his reputation as a pioneer of modern experimental science by publishing New Experiments and Observations Touching Cold. During this period, Boyle and his contemporaries, including Robert Hooke, emphasize that chemical understanding must be grounded in controlled experimentation, precise documentation of apparatus, and the replication of results.[38] England
1700 Concept Korean mathematician Choi Seok-jeong is the first to publish an example of Latin squares of order nine, in order to construct a magic square, predating Leonhard Euler by 67 years.[39] Latin squares are used in combinatorics and in experimental design.[40] Korea
1710 Concept John Arbuthnot publishes "An Argument for Divine Providence, taken from the constant regularity observed in the births of both sexes" in Philosophical Transactions of the Royal Society, examining London birth records for each of 82 years from 1629 to 1710 and finding that male births exceeded female births in every single year. Applying what becomes known as the sign test, Arbuthnot calculates that the probability of this occurring by chance alone (if male and female births were equally likely) is 0.582 — an astronomically small figure — and concludes the pattern reflects design rather than chance. The paper is widely credited as the earliest recorded use of statistical hypothesis testing, predating the formal apparatus of significance testing by over two centuries.[41] England
1713 Concept Jacob Bernoulli formulates the law of large numbers, establishing convergence of sample averages to expected values. Switzerland
1747 Experiment Scottish doctor James Lind conducts one of the earliest controlled clinical trials when investigating the efficacy of citrus fruit in cases of scurvy, dividing twelve scurvy patients, whose "cases were as similar as I could have them", into six pairs — without randomization, but with an early attempt at controlling for similar starting conditions. Each pair is given a different remedy. According to Lind's 1753 Treatise on the Scurvy in Three Parts Containing an Inquiry into the Nature, Causes, and Cure of the Disease, Together with a Critical and Chronological View of what has been Published of the Subject, the remedies were: one quart of cider per day, twenty-five drops of elixir vitriol (sulfuric acid) three times a day, two spoonfuls of vinegar three times a day, a course of sea-water (half a pint every day), two oranges and one lemon each day, and electuary, (a mixture containing garlic, mustard, balsam of Peru, and myrrh).[42][22] Lind would note that the pair who had been given the oranges and lemons were so restored to health within six days of treatment that one of them returned to duty, and the other was well enough to attend the rest of the sick.[42] United Kingdom
1778 Concept Pierre-Simon Laplace examines birth statistics from almost half a million births in "Mémoire sur les probabilités," finding a consistent excess of male births over female births and concluding through probabilistic calculation that the excess is a real, though unexplained, effect rather than a product of chance — extending the sex-ratio hypothesis-testing tradition begun by John Arbuthnot in 1710 to a substantially larger dataset and a more developed probabilistic apparatus.[43] France
1782 Concept Leonhard Euler publishes "Recherches sur une nouvelle espèce de quarrés magiques" in the Verhandelingen of the Zeeland Society of Sciences (first presented to the Imperial Academy of Sciences of St. Petersburg in 1779), using Latin characters as symbols in his arrays and thereby giving rise to the name "Latin square." Though Korean mathematician Choi Seok-jeong had published an example of order-nine Latin squares 67 years earlier, Euler's work begins the general mathematical theory of Latin squares and their combinatorial properties.[44] Russia
1780 Concept Franz Mesmer, then already well known in Paris, refuses a request from the Paris Faculty of Medicine's Joseph-Marie-François de Lassone to have his therapeutic method scrutinized through new cures on unfamiliar patients, arguing instead that his prior "cures" should be taken as an objective matter of record — a position that both 1784 Royal Commissions would later reject in designing their own investigations. Separately, Mesmer proposed principles for objectively testing his method that anticipate elements of controlled experimental design.[45] France
1784 Experiment The first blinded experiment is conducted by the French Academy of Sciences to investigate the claims of mesmerism as proposed by Franz Mesmer. In the experiment, researchers blindfolded mesmerists and asked them to identify objects that the experimenters had previously filled with "vital fluid". The subjects are unable to do so.[46] France
1784 Experiment Two independent French Royal Commissions — a nine-member "Franklin Commission" (four physicians from the Paris Faculty of Medicine and five scientists from the Royal Academy of Sciences, including Benjamin Franklin and Antoine Lavoisier) and a five-member "Society Commission" from the Royal Society of Medicine — are appointed by Louis XVI to investigate Charles d'Eslon's claims for the physical existence of "animal magnetism." Lavoisier designs the Franklin Commission's investigative protocol: rather than assessing long-term cures, the commissioners test the "momentary" physiological effects of magnetization under conditions systematically varying whether subjects are genuinely or falsely told they are being magnetized, and whether investigators and subjects are blindfolded. In one representative test, a "sensitive" subject who believed he had been led to a magnetized tree fainted at the foot of the wrong tree; in another, a subject who drank ordinary water believing it magnetized displayed magnetic symptoms. Both Commissions conclude that d'Eslon's magnetic fluid does not exist and that all observed effects are attributable to touch, imagination, and imitation. The Franklin Commission's report, presented 11 August 1784 and printed in at least 20,000 copies, is widely regarded by historians of science as one of the earliest classic examples of a systematic controlled trial, notable in particular for its use of literal blindfolding of both investigators and subjects.[47][48] France
1796 Experiment English physician Edward Jenner vaccinates James Phipps, an eight-year-old boy, with cowpox material as a test of protection against smallpox — a landmark early medical experiment, though, like Lady Mary Wortley Montagu's earlier variolation trials on condemned prisoners, conceptually flawed by the complete absence of a scientific control to establish whether the observed protection was actually due to the intervention.[49] England
1798 Statistical method German mathematician Carl Friedrich Gauss develops the mathematical foundations of the method of least squares. Years later, the method would enable astronomers to predict the orbit of the asteroid Ceres after its discovery by Giuseppe Piazzi and help Franz Xaver von Zach successfully relocate it.[50] Germany
1799 Concept English physician John Haygarth demonstrates the importance of a control group for correctly identifying the placebo effect, in his study of "Perkin's tractors" — a popular but ineffective metallic device claimed to cure disease — by comparing its effects against sham wooden tractors, finding no difference in outcome between the real device and the placebo.[51] England
1815 Concept An article on optimal designs for polynomial regression is published by Joseph Diaz Gergonne.[52] France
1817 Experiment The first blinded experiment recorded outside of a scientific setting compares the musical quality of a Stradivarius violin to one with a guitar-like design. A violinist plays each instrument while a committee of scientists and musicians listen from another room so as to avoid prejudice.[53][54] France
1827 Method Pierre-Simon Laplace uses least squares methods to address analysis of variance problems regarding measurements of atmospheric tides.[55] France
1835 Experiment An early example of a double-blind protocol is the Nuremberg salt test performed by Friedrich Wilhelm von Hoven, Nuremberg's highest-ranking public health official.[56] Germany
1835 Concept Belgian statistician Adolphe Quetelet introduces the concept of the “average man” (l’homme moyen), arguing that human physical and social traits follow statistical distributions. By applying probability and quantitative analysis to crime, mortality, and demographics, he helps establish statistical approaches in sociology, demography, and the emerging social sciences.[57] Belgium
1843 Concept Irish doctor James Henry proposes principles for conducting a controlled trial comparing cold-water therapy against sulphur treatment for scabies — an early articulation of controlled-trial methodology predating the formalization of randomized experimental design by nearly a century.[58] Ireland
1844 Concept Wesley S. B. Woolhouse poses the problem of what becomes known as Steiner triple systems as Prize Question #1733 in the Lady's and Gentlemen's Diary — asking for a systematic partition of elements into triples such that every pair of elements appears together in exactly one triple.[59] England
1847 Method Thomas Kirkman solves Wesley S. B. Woolhouse's 1844 prize problem in "On a Problem in Combinations," published in The Cambridge and Dublin Mathematical Journal — the first solution to what becomes known as the Steiner triple system problem, three years before Kirkman posed his own, more elaborate variant asking for resolvable systems (later known as Kirkman's schoolgirl problem).[60] England
1850 Concept English clergyman and mathematician Thomas Kirkman poses what becomes known as Kirkman's schoolgirl problem: whether fifteen schoolgirls can walk in five groups of three every day for seven days such that no two girls walk in the same group more than once. A solution to this recreational problem is equivalent to a resolution of a balanced incomplete block design with 15 points, block size 3, and λ = 1 — an early landmark in the theory of resolvable designs.[61] England
1853 Concept Jakob Steiner independently reintroduces triple systems in "Combinatorische Aufgabe," published in Crelle's Journal, unaware of Thomas Kirkman's 1847 solution to the same problem. Because Steiner's paper becomes more widely known within the mathematical community than Kirkman's earlier work, the systems come to be named Steiner systems in his honor rather than Kirkman's, despite Kirkman's priority.[62] Germany
1854 Experiment English physician John Snow investigates a severe cholera outbreak in London by mapping cholera deaths geographically, identifying a shared water pump on Broad Street as the likely common source of infection — considered the first ecological study to solve a public health problem, using population-level (group) data rather than individual-level observation. Snow has the pump handle removed, after which deaths in the area subside. The underlying mechanism of cholera transmission by bacteria would not be understood until Robert Koch's later discoveries.[63] England
1860 Method German psychologist and physicist Gustav Theodor Fechner makes a groundbreaking contribution to the development of experimental psychology with the publication of Elements of Psychophysics. In this work, he seeks to test and justify the relationship between physical stimuli and the sensations they produce. Fechner proposes that mental experiences can be quantified by linking them to measurable physical changes, laying the foundation for psychophysics. Using experimental data, he formulates mathematical laws—most notably the Weber-Fechner law—that describe how perceived intensity varies with stimulus magnitude. His efforts help establish psychology as a quantitative science rooted in empirical observation.[64] Germany
1861 Experiment French chemist and microbiologist Louis Pasteur conducts controlled experiments demonstrating that microorganisms originate from airborne particles rather than spontaneous generation. In his memoir examining this doctrine, he shows that sterilized broth remain free of life unless exposed to contaminated air, helping establish the foundations of modern microbiology and germ theory.[65][66] France
1865 Concept French physiologist Claude Bernard publishes Introduction to the Study of Experimental Medicine, advocating that researchers should not know the hypothesis being tested while making observations — a recommendation that directly contradicted the prevailing Enlightenment-era view that valid scientific observation required a well-educated, informed observer.[67] France
1876 Concept American scientist Charles S. Peirce contributes the first English-language publication on an optimal design for regression models.[68] United States
1877 Concept Charles Sanders Peirce formalizes inquiry as a structured experimental process: hypotheses (abduction) generate testable predictions (deduction), which are evaluated through observation and experiment (induction). This integrates reasoning with empirical testing, establishing a cyclical, self-correcting model of experiment design focused on hypothesis testing and iterative refinement, developed across his "Illustrations of the Logic of Science" series.[69][70] United States
1879 Organization Experimental psychology emerges as a modern scientific discipline in Germany with the establishment of the first experimental laboratory by Wilhelm Wundt at the University of Leipzig. This marks a pivotal moment in the history of psychology, as Wundt seeks to separate psychology from philosophy by applying rigorous scientific methods. He introduces a mathematical and experimental approach to studying the human mind, emphasizing observation, measurement, and controlled experimentation. Wundt's work lays the foundation for psychology as an empirical science, influencing future researchers and schools of thought.[64] Germany
1882 Concept In his published lecture at Johns Hopkins University, Peirce introduces experimental design with these words:

Logic will not undertake to inform you what kind of experiments you ought to make in order best to determine the acceleration of gravity, or the value of the Ohm; but it will tell you how to proceed to form a plan of experimentation.

[....] Unfortunately practice generally precedes theory, and it is the usual fate of mankind to get things done in some boggling way first, and find out afterward how they could have been done much more easily and perfectly.[71]

United States
1884 Organization Frederick Akbar Mahomed, who worked at Guy's Hospital in London and separated chronic nephritis with secondary hypertension from what would later be termed essential hypertension, founds the Collective Investigation Record for the British Medical Association — an organization collecting data from physicians practicing outside the hospital setting, considered the precursor of modern collaborative, multi-site clinical trials.[72] England
1885 Concept Analysis of variance. An eloquent non-mathematical explanation of the additive effects model becomes available.[73] United Kingdom
1887 Experiment Albert A. Michelson and Edward W. Morley conduct an interferometry experiment designed to detect Earth's motion through the hypothesized luminiferous aether by measuring expected variations in the speed of light in different directions. The experiment fails to detect the predicted aether drift, instead measuring a drift far too small to account for the theoretically expected effect and generally attributed to noise — becoming one of the most famous null results in the history of science. The unexpected absence of a positive result would go on to contribute significantly to the development of special relativity.[74] United States
1880s Experiment Charles Sanders Peirce and Joseph Jastrow introduce randomized experiments in the field of psychology.[75] United States
1896 Concept French psychologist Alfred Binet presents early scientific work on human perception of magic (illusion) tricks, examining how magicians exploit blind spots in attention and perception — among the earliest documented scientific interest in magic as a subject of study, though the field would not be substantially revived until the early 21st century.[76] France
1897 Experiment Norman Triplett conducts one of the first social psychology experiments on cyclist performance.[77] United States
1900 Method The P-value is first formally introduced by Karl Pearson, in his Pearson's chi-squared test, using the chi-squared distribution and notated as capital P.[78] Since then, P-values would become the preferred method to summarize the results of medical articles.[79][80] United Kingdom
1900 Concept French mathematician Gaston Tarry proves there is no pair of orthogonal Latin squares of order six, resolving Leonhard Euler's 36 officers problem in the negative. The result is later recognized as equivalent to the nonexistence of a (43,7,1) balanced incomplete block design — an early demonstration that not every combinatorially plausible block-design parameter set is actually realizable.[81] France
1901 Concept The principle of informed consent is first enacted, without general or official guidance, in the United States Army's research into yellow fever transmission in Cuba — an early instance of volunteers being informed of and consenting to the risks of an experiment before participation. The precedent would later be referenced during the drafting of the Nuremberg Code decades later.[82] Cuba
1903 Concept American physician Richard Clarke Cabot concludes that the placebo should be avoided because it is deceptive.[83] United States
1907 Experiment German philosopher and psychologist Carl Stumpf, together with his assistant Oskar Pfungst, investigates the claims surrounding "Clever Hans," an Orlov Trotter horse whose owner, Wilhelm von Osten, claimed could perform arithmetic and other cognitive tasks. Ruling out deliberate fraud, Pfungst determines the horse answers correctly 89% of the time when the questioner knows the answer, but only 6% of the time when the questioner does not — and shows that questioners unconsciously tensed their posture and expression as the horse's taps approached the correct answer, providing an unintentional cue the horse had learned to use as a stopping signal. The episode becomes the classic illustration of the observer-expectancy effect and a foundational argument for blinded experiment the questioner or experimenter, not just the subject, from knowledge that could unconsciously bias results.[84] Germany
1907 Experiment The first study recorded to have a blinded researcher is conducted by W. H. R. Rivers and H. N. Webber to investigate the effects of caffeine.[85] United Kingdom
1908 Method British statistician William Sealy Gosset, working at Guinness, introduces the Student’s t-distribution to analyze small sample data when population variance is unknown. Publishing under the pseudonym “Student” in the journal Biometrika, he provides a statistical method widely used for inference with limited data.[86] United Kingdom
1908–1909 Concept Irish economist and statistician Francis Ysidro Edgeworth publishes a series of papers "On the Probable Errors of Frequency-Constants" that anticipate several elements of what would become known as Fisher information, over a decade before Fisher's own 1922 formal treatment — work later recognized by statisticians and historians of statistics as an important, underacknowledged precursor.[87] Ireland
1918 Concept English statistician Ronald Fisher introduces the term variance and proposes its formal analysis in his article The Correlation Between Relatives on the Supposition of Mendelian Inheritance.[88] United Kingdom
1918 Method Kirstine Smith publishes a major study in Biometrika analyzing the statistical precision of polynomial regression estimates. She derives principles for choosing observation points that minimize estimation variance, introducing foundational methods for optimal experimental design in polynomial models and influencing later developments in statistical design theory.[89] Denmark
1918–1940s Method Ronald A. Fisher and collaborators establish the foundations of modern experimental design in agricultural research, introducing factorial designs and analysis of variance (ANOVA).[14] United Kingdom
1919 Organization Ronald A. Fisher is hired as a statistician at the Rothamsted Experimental Station, where his work on poorly designed agricultural data helps trigger the modern statistical approach to experimental design.[19] United Kingdom
1919 Method R. A. Fisher at the Rothamsted Experimental Station in England starts developing modern concepts of experimental design in the planning of agricultural field experiments.[90] England
1921 Method Ronald Fisher publishes his first application of the analysis of variance.[91] United Kingdom
1922 Concept Ronald Fisher formally defines what becomes known as Fisher information in "On the Mathematical Foundations of Theoretical Statistics," measuring the amount of information an observable random variable carries about an unknown parameter of the distribution generating it. Fisher information becomes foundational to the theory of maximum-likelihood estimation and to optimal design of experiments, where maximizing the Fisher information of a design corresponds to minimizing the variance of parameter estimates obtainable from it.[92] United Kingdom
1923 Field development The first randomization model is published in Polish by Jerzy Neyman.[93] Poland
1923 Method Ronald A. Fisher, together with Winifred Mackenzie, publishes "Studies in Crop Variation II," extending his 1921 "Studies in Crop Variation I" (which had partitioned the variation of a time series into annual and slow-deterioration components) to study yield variation across plots sown with different potato varieties and subjected to different fertiliser treatments — early applied work at Rothamsted Experimental Station that would feed directly into the formal analysis-of-variance framework Fisher developed over the following years.[94][95] United Kingdom
2024 Concept Cochrane (organisation) publishes an updated meta-epidemiological review by Ingrid Toews and colleagues comparing healthcare outcomes assessed in observational studies against those assessed in randomized controlled trials, finding little evidence of significant effect differences between the two designs regardless of the specific observational design used — a notably more optimistic finding for observational methods than the experimental-benchmarking critiques raised by Robert LaLonde (1986) and others in economics and social science.[96] Global
1924–1925 Experiment Political scientist Harold Gosnell conducts an early field experiment on voter participation in Chicago, testing whether nonpartisan mailings encouraging registration and voting increased turnout — widely regarded as one of the earliest field experiments in political science, using randomization outside a laboratory setting to study a real-world civic behavior.[97] United States
1925 Publication British statistician Ronald Fisher publishes Statistical Methods for Research Workers, developing tests of statistical significance, proposing the 0.05 threshold for hypothesis testing, and popularizing analysis of variance as a method for partitioning variation in experimental data. The book would later be identified as a seminal, and controversial, influence — some retrospective critics tracing aspects of the modern replication crisis in science to conventions it popularized.[98][99][100][101][102][24][103] United Kingdom
2025 Policy Miguel Hernán, Aidan G. Cashin, and an international group of collaborators publish the TARGET Statement in JAMA, a reporting guideline for observational studies that use the "target trial emulation" framework — designing an observational study to explicitly emulate the randomized controlled trial that would ideally answer the same causal question, in order to reduce common biases in observational causal inference. The statement extends the lineage of trial-reporting guidelines established by CONSORT (for randomized trials) and STROBE (for observational studies generally) to this specific and increasingly widely used design.[104] United States
1926 Publication John Russell (agricultural scientist) publishes "Field Experiments: How They Are Made and What They Are", summarizing contemporary practices and principles of experimental design in agricultural research.[105] United Kingdom
1926 Concept Ronald Fisher argues in "The Arrangement of Field Experiments" that studying multiple factors simultaneously in "complex" (factorial) designs is more efficient than the traditional one-factor-at-a-time approach, writing that "Nature... will best respond to a logical and carefully thought out questionnaire; indeed, if we ask her a single question, she will often refuse to answer until some other topic has been discussed" — arguing that a well-designed factorial experiment can determine the effects of several factors, and their interactions, using no more trials than would be needed to determine just one factor's effect alone.[106] United Kingdom
1926 Experiment Janet Lane-Claypon publishes a retrospective investigation of risk factors for breast cancer, widely regarded as one of the first major case–control studies — comparing women with the disease against a control group without it to examine differences in exposure to suspected risk factors.[107] United Kingdom
1926 Concept Soviet mathematician Sergei Natanovich Bernstein introduces the "blocks method" in probability theory, splitting a sample into blocks separated by smaller subblocks so the blocks can be treated as approximately independent — a technique later used to prove limit theorems for sums of dependent random variables and applied in extreme value theory. This is a distinct lineage from Fisher's contemporaneous statistical blocking in experimental design, converging on the same term for a different underlying idea.[108] Soviet Union
1933 Concept Jerzy Neyman and Egon Pearson publish "On the Problem of the Most Efficient Tests of Statistical Hypotheses" in Philosophical Transactions of the Royal Society A, part of a series of papers begun in 1928 that formalizes the statistical hypothesis test as a proposed refinement of Ronald Fisher's significance-testing approach. The papers supply much of the standard terminology still used today, including the term "alternative hypothesis" and the H0/H1 notation for the null and alternative hypotheses, and also introduce the simple/composite hypothesis distinction. Fisher and Neyman would go on to quarrel over the relative merits of their competing formulations until Fisher's death in 1962; the two approaches were later merged into a single hybrid framework by textbook writers and practitioners without direct input from either originator.[109] United Kingdom
1933 Method English mathematician Raymond Paley publishes a construction method for generating orthogonal matrices whose elements are all either +1 or −1 (Hadamard matrix), working for matrices of size N for most N equal to a multiple of 4 (all such N up to 100 except N = 92). The Paley construction becomes the mathematical foundation Robin Plackett and J. P. Burman rely on in 1946 to generate their two-level factorial screening designs.[110] United Kingdom
1935 Method Jerzy Neyman formulates a randomization model for randomized block designs, extending his 1923 randomization model for completely randomized designs.[111] Poland
1935 Concept Ronald Fisher introduces the term "confounding" in The Design of Experiments to describe a specific consequence of blocking (statistics) in a factorial experiment — when partitioning treatment combinations into blocks causes certain interaction effects to become indistinguishable from ("confounded with") block effects. Fisher's usage popularizes the term within statistics, though his concern was controlling heterogeneity among experimental units rather than causal inference in the modern sense.[112] United Kingdom
1935 Publication Ronald Fisher publishes The Design of Experiments, a foundational work that formalizes principles such as randomization, replication, and blocking, and emphasizes the importance of efficient experimental design; the book becomes the basis of modern experimental science and remains one of the field's foundational texts.[113][114][115][24][22][20] United Kingdom
1935 Experiment Ronald Fisher describes the "lady tasting tea" experiment in The Design of Experiments, the original exposition of his concept of the null hypothesis — which he describes as never proved or established, only possibly disproved through experimentation. The example is loosely based on a real event: phycologist Muriel Bristol claimed she could tell whether tea or milk was poured into a cup first, and her future husband William Roach suggested testing her with eight cups, four of each preparation, in random order. Fisher's treatment, using what becomes known as Fisher's exact test, calculates that correctly identifying all eight cups would occur by chance alone in only 1 of 70 possible combinations (about 1.4%), below the conventional 5% significance threshold — establishing the logic of randomized assignment combined with a combinatorial significance test as a template later applied throughout experimental science. According to statistician H. Fairfield Smith, as later reported by David Salsburg, Bristol did in fact identify all eight cups correctly in the real trial.[116][117] United Kingdom
1935 Concept The term "factorial" (in the experimental-design sense) appears to enter print for the first time when Ronald Fisher uses it in The Design of Experiments.[118] United Kingdom
1937 Method Researcher Meredith Crawford, at the Yerkes National Primate Research Center, invents the cooperative pulling paradigm — an experimental design in which two or more animals must pull a reward toward themselves via an apparatus neither can operate alone — publishing a study of two young chimpanzees, Bula and Bimba, coordinating to pull ropes attached to a box too heavy for either to move alone. The paradigm becomes the most widely used experimental design for testing cooperation in animals, later applied to bonobos, orangutans, capuchins, elephants, wolves, dogs, ravens, kea, and dolphins, among other species, to investigate the evolution and mechanisms of cooperative behavior.[119] United States
1939 Concept A publication by Bose and Nair underlie the concept of association scheme. In their paper, they introduced the concept of association schemes as a way to study the structure of contingency tables. They show that association schemes can be used to represent the dependencies between the variables in a contingency table, and that they can be used to derive statistical tests for independence.[120] India
1940 Method Ronald A. Fisher publishes "An Examination of the Different Possible Solutions of a Problem in Incomplete Blocks" in Annals of Eugenics, extending his earlier work on blocking (statistics) to systematically examine balanced incomplete block designs — designs in which not every treatment can be tested in every block, a situation Fisher's original randomized block design did not address.[121] United Kingdom
1940 Method Raj Chandra Bose and K. Kishen at the Indian Statistical Institute independently find some efficient designs for estimating several main effects. India
1942 Concept K. Kishen generalizes Latin squares and mutually orthogonal Latin squares to Latin cubes and Latin hypercubes in "On latin and hyper-graeco cubes and hypercubes," published in Current Science.[122] India
1943 Method Economist Robert Dorfman introduces group testing in a short report published in the Annals of Mathematical Statistics, motivated by the United States Public Health Service's wartime effort to screen soldiers for syphilis without the expense of testing every individual blood sample separately. Dorfman's method pools blood samples into groups and tests each pooled sample; groups testing negative allow every soldier within them to be cleared with a single test, while positive groups require individual follow-up testing — dramatically reducing the expected number of tests needed when disease prevalence is low. Dorfman tabulates the optimal group size as a function of prevalence rate.Cite error: Closing </ref> missing for <ref> tag[123] United Kingdom
1945 Method American statistician Abraham Wald publishes Sequential Tests of Statistical Hypotheses, pioneering the field of sequential analysis — a framework for statistical tests where the number of observations is not fixed in advance, and the decision to stop collecting data (accept, reject, or continue sampling) can depend on the accumulated results as the experiment proceeds. Wald's work, developed partly in a wartime context, provides the formal apparatus later applied throughout industrial and clinical sequential experimental design.[124] United States
1945 Concept British statistician D. J. Finney publishes "The Fractional Replication of Factorial Arrangements" in Annals of Eugenics, containing the first statistical use in print of the term "aliasing (factorial experiments)" — the phenomenon in fractional factorial designs where some effects become indistinguishable from each other because only a fraction of all possible treatment combinations is observed. The term would later be adopted into signal processing theory, possibly influenced by this earlier statistical usage.[125] United Kingdom
1945 Method British statistician D. J. Finney introduces fractional factorial design, extending Ronald Fisher's earlier work on full factorial experiments at the Rothamsted Experimental Station by showing how to test only a carefully chosen fraction of all possible factor-level combinations — reducing the number of experimental runs required while deliberately confounding selected effects, based on the assumption that higher-order interactions are typically negligible (the sparsity-of-effects principle). Originally developed for agricultural applications, fractional factorial design later spreads to engineering, business, and other sciences.[126] United Kingdom
1946 Experiment Psychologist E.M. Jellinek conducts a crossover trial for a U.S. headache-remedy manufacturer testing whether removing a scarce wartime ingredient would reduce drug efficacy, randomly assigning 199 subjects with frequent headaches to rotate through four formulations (three real combinations of ingredients plus a lactate placebo) over eight weeks. Initial analysis across all subjects suggests the scarce ingredient is unnecessary, but Jellinek discovers that 120 of the 199 subjects are "placebo reactors" whose inclusion masks the ingredient's real effect; restricting analysis to the 79 non-reactors reveals the ingredient does contribute significantly to efficacy. The trial becomes an influential early demonstration that placebo responsiveness varies systematically across individuals and can confound drug-efficacy comparisons if not accounted for.[127] United States
1946 Method R.L. Plackett and J.P. Burman publish a renowned paper titled "The Design of Optimal Multifactorial Experiments". The paper introduces what would be called Plackett–Burman designs, which are highly efficient screening designs with run numbers that are multiples of four. These designs are particularly useful for experiments where only main effects are of interest. In a Plackett-Burman design, main effects are often heavily confounded with two-factor interactions, making them suitable for screening experiments. For instance, a Plackett-Burman design with 12 runs can be utilized for an experiment containing up to 11 factors.[128] United Kingdom
1946–1947 Concept C. R. Rao generalizes Kishen's hypercube constructions to arrays of arbitrary strength t in "Hypercubes of strength d leading to confounded designs in factorial experiments" (1946), then introduces the general concept of the orthogonal array — a tabular structure in which every selection of t columns contains all possible t-tuples of symbols the same number of times — in "Factorial experiments derivable from combinatorial arrangements of arrays" (1947), the paper credited as the origin of the modern notion of orthogonal arrays used throughout combinatorial design theory, coding theory, cryptography, and software testing.[129][130] India
1946–1947 Experiment Sir Geoffrey Marshall of the MRC Tuberculosis Research Unit conducts the first randomized curative trial, testing the efficacy of streptomycin for pulmonary tuberculosis — both double-blind and placebo-controlled — the same landmark trial elsewhere credited to the Medical Research Council (United Kingdom) as an institution, here attributed to its individual lead investigator.[131] United Kingdom
1947 Policy The Nuremberg Code is formulated as a result of the Doctors' Trial at the Nuremberg trials, in which Nazi physicians were tried for murdering and torturing concentration camp prisoners in valueless medical experiments — several were subsequently hanged. The Code establishes foundational ethical principles for human experimentation, including that experimenters should not subject participants to procedures they would not undertake themselves, and becomes the first codification of research ethics standards, influencing medical experiment codes of practice worldwide.[132] Germany
1947 Policy The Nuremberg Code is articulated by the U.S. military tribunal in USA v. Brandt (the "Doctors' Trial"), part of the Subsequent Nuremberg Trials against German physicians who conducted unethical experiments on concentration-camp prisoners and carried out over 3.5 million forced sterilizations. Concerned that the defendants — who argued their experiments differed little from prewar research and that no law distinguished legal from illegal experimentation — might escape conviction, prosecution medical experts Andrew Conway Ivy and Leo Alexander drafted a memorandum in April 1947 outlining principles for legitimate medical research; an expanded version requiring explicit voluntary consent followed on 9 August 1947. The judges' verdict, delivered 20 August 1947 against Karl Brandt (physician) and 22 others, revised these into ten points establishing principles including informed consent, freedom from coercion, and beneficence toward research subjects — the first codification of research ethics standards, influencing medical experiment codes of practice worldwide despite going largely unenforced (and even dismissed by some as inapplicable to "ordinary physicians") for roughly two decades after it was written.[133][134] Germany
1948 Concept Francis J. Anscombe, at Rothamsted Experimental Station, discusses and develops design-based (randomization-based) analysis of variance, later extended by Oscar Kempthorne at Iowa State University, who introduces the assumption of unit-treatment additivity — the idea that an experimental unit's observed response can be written as the sum of the unit's baseline response and a treatment effect that is constant across all units receiving that treatment. This randomization-based approach differs from the more commonly taught normal-linear-model approach in making no assumption of a normal distribution or of independence between observations, though the two approaches' test statistics closely approximate each other in practice.[135] United Kingdom
1948 Experiment The Medical Research Council conducts a landmark randomized controlled trial of streptomycin for pulmonary tuberculosis, establishing modern clinical trial methods such as random allocation and blinding. Results show reduced mortality but emerging drug resistance and side effects, highlighting need for combination therapies and shaping evidence-based medicine future practice.[136] United Kingdom
1948 Concept British statistician Frank Yates introduces the concept of restricted randomization.[137][138] United Kingdom
1949 Method Psychologist Richard Solomon (psychologist) develops the Solomon four-group design, a research method addressing the problem of pretest sensitization — the possibility that administering a pre-intervention test itself influences how subjects respond to a subsequent treatment. In addition to the standard pre-test/treatment/post-test and pre-test/control/post-test groups, the design adds two further groups that skip the pre-test entirely (treatment-only and control-only, each still post-tested), allowing researchers to separate the effect of the treatment itself from the effect of having taken the pre-test.[139] United States
1949 Method Genichi Taguchi begins developing his experimental design techniques while working at Japan’s Electrical Communications Laboratories (ECL) in the post–World War II reconstruction period. Tasked with improving research and development productivity, he formulates a new approach to quality improvement that emphasizes off-line quality control, robust design, and statistical experimentation. These early efforts lay the foundations of what would later become known as the Taguchi Methods, integrating engineering design with statistical principles to systematically reduce variability, improve product quality, and lower societal and manufacturing costs.[140] Japan
1949 Method Kenneth Arrow, David Blackwell, and M.A. Girshick publish "Bayes and Minimax Solutions of Sequential Decision Problems" in Econometrica, an early and influential contribution to sequential decision theory building on Abraham Wald's wartime work on sequential analysis, formalizing Bayes and minimax approaches to problems where the decision of when to stop sampling is itself part of the optimization.[141] United States
1949 Concept R. H. Bruck and H. J. Ryser prove a nonexistence result for finite projective planes — special cases of symmetric block designs with λ = 1 — showing that if a projective plane of order q exists and q ≡ 1 or 2 (mod 4), then q must be expressible as the sum of two squares. The result rules out projective planes of certain orders (such as 6) that satisfy the basic combinatorial parameter equations for a symmetric design but are nonetheless impossible.[142] United States
1950 Publication Gertrude Mary Cox and William Gemmell Cochran publish the book Experimental Designs, which would become the major reference work on the design of experiments for statisticians for years afterwards.[143] United States
1950 Method P.M. Grundy and Michael Healy (statistician) publish "Restricted Randomization and Quasi-Latin Squares" in the Journal of the Royal Statistical Society, Series B, developing restricted randomization — a method for excluding intuitively poor or undesirable treatment allocations from the space of possible randomizations (for example, preventing a new obesity treatment from being randomly assigned only to the heaviest patients) while still preserving the theoretical statistical benefits of randomization. The concept had also been introduced independently by Frank Yates, whose earlier work on the subject this timeline's 1948 entry describes.[144] United Kingdom
1950 Concept K. Bush coins the term "orthogonal array" in his PhD thesis at the University of North Carolina, naming the structure C. R. Rao had introduced in 1947 (Rao himself had used the unmodified term "array," meaning simply a subset of treatment combinations, before the tabular/matrix framing made the possibility of non-simple arrays — those with repeated rows — apparent).[145] United States
1950 Method Cochran and Cox formalize principles of sample size determination in experimental design, providing more accurate tables and methods.[20] United States
1950 Experiment Richard Doll and Austin Bradford Hill publish a preliminary report demonstrating a statistically significant association between tobacco smoking and lung cancer, using a large case–control study design comparing lung cancer patients against controls without the disease. Critics initially argued the case–control design could not establish causation; subsequent cohort studies by the same researchers over following decades confirmed the causal link the case–control study had suggested, and smoking is now recognized as the cause of roughly 87% of lung cancer deaths in the United States.[146] United Kingdom
1950 Concept S. Chowla and H. J. Ryser extend the 1949 Bruck–Ryser nonexistence result from projective planes to general symmetric block designs, establishing what becomes known as the Bruck–Ryser–Chowla theorem: necessary conditions (a square-difference condition when the number of points is even, and a solvable Diophantine equation when odd) that any symmetric (v, b, r, k, λ)-design's parameters must satisfy, ruling out entire classes of otherwise combinatorially plausible designs.[147] United States
1950 Concept Randomized controlled trials begin to emerge as the gold standard in medical research, enabling systematic causal inference Global
1951 Method Formal placebo-controlled and multi-arm clinical trial designs are developed to improve reliability of treatment comparisons United States
1951 Method George E. P. Box and K. B. Wilson introduce response surface methodology (RSM), using sequences of designed experiments and second-degree polynomial models to approximate and optimize responses, even with limited process knowledge. They develop a method to build a quadratic model where the number of data points scales linearly, rather than exponentially, with the number of inputs, striking a balance between accuracy and ease of application — enabling sequential, adaptive optimization of industrial processes even under limited prior knowledge of the underlying system.[148][149][14][24][21] United Kingdom
1952 Concept American mathematician and statistician Herbert Robbins recognizes the significance of a problem where a gambler faces a trade-off between "exploitation" of the machine with the highest expected payoff and "exploration" to learn about other machines' payoffs. This problem involves pulling levers on different machines, each providing random rewards from unknown probability distributions. The gambler aims to maximize the total rewards earned over a sequence of lever pulls. Robbins devised convergent population selection strategies in his work on "some aspects of the sequential design of experiments."[150] United States
1952 Method D.G. Horvitz and D.J. Thompson publish "A Generalization of Sampling Without Replacement from a Finite Universe" in the Journal of the American Statistical Association, introducing what becomes known as the Horvitz–Thompson estimator — a method for producing unbiased population estimates from samples with unequal selection probabilities by weighting each observation by the inverse of its probability of inclusion. Originally developed for survey sampling, the estimator later becomes a standard tool via inverse probability weighting for correcting bias in the analysis of experiments with unequal exposure probabilities, including studies of spillover and network interference effects.[151] United States
1952 Concept development Bose and Shimamoto introduce the term association scheme.[152] United States
1954 Experiment The Salk polio vaccine trial becomes one of the first large-scale randomized controlled trials, involving over one million participants United States
1954 Concept Paul Lazarsfeld and others formalize factor analysis as a tool for latent variable modeling in experiments. United States
1954 Publication Edwin Boring discusses the history and meanings of control treatments, clarifying the role of controls in experimental research.[20] United States
1954 Publication American experimental psychologist Edward Boring writes an article titled The History of Experimental Design. In this article, Boring notes that the early history of ideas on the planning of experiments has been "but little studied".[90] United States
1955 Experiment An influential study entitled The Powerful Placebo firmly establishes the idea that placebo effects are clinically important.[153] United States
1955 Method M. B. Wilk introduces the randomization analysis of the generalized randomized block design (GRBD), which replicates each treatment at least twice within each block — unlike the classic randomized block design, which has no within-block replication — allowing block-treatment interaction (statistics) to be estimated and tested without relying on parametric assumptions about the distribution of experimental error.[154] United States
1956 Concept D. V. Lindley publishes "On a Measure of Information Provided by an Experiment," proposing that the utility of an experimental design be measured as the Kullback–Leibler divergence between the posterior and prior probability distributions over the parameters being estimated — showing this expected utility is coordinate-independent and equals the mutual information between the parameters and the observations, exactly the expected information gain from running the experiment. This becomes a foundational formulation within Bayesian experimental design.[155] United Kingdom
1956 Experiment Richard Doll and Austin Bradford Hill publish a second report on the mortality of British doctors, a prospective cohort study confirming the causal link between smoking and lung cancer that their 1950 case–control study had first suggested — an early demonstration of how case–control findings could be validated through a stronger, prospective study design.[156] United Kingdom
1959 Concept Leslie Kish uses the term "confounding" in the sense of "incomparability" between two or more groups (such as exposed and unexposed groups) in an observational study — a distinct usage from Fisher's blocking-based sense, and closer to the modern causal-inference meaning of the term as it would later develop in epidemiology.[157] United States
1959 Method Raymond H. Pierson and Edward A. Fay publish "Guidelines for Interlaboratory Testing Programs" in Analytical Chemistry, an early formalization of round-robin testing — interlaboratory studies in which multiple independent laboratories perform the same measurement or analysis, using the same or different methods and equipment, to assess the reproducibility of a test method or verify a new method of analysis against an established one.[158] United States
1959 Concept Raj Chandra Bose and D. M. Mesner publish "On Linear Associative Algebras Corresponding to Association Schemes of Partially Balanced Designs" in the Annals of Mathematical Statistics, introducing what becomes known as the Bose–Mesner algebra — the commutative algebra generated by the adjacency matrices of an association scheme's relations. This publication turns association schemes, previously a statistical tool for classifying partially balanced incomplete block designs, into an object of independent algebraic interest.[159] United States
1959 Concept W. M. S. Russell and R. L. Burch publish The Principles of Humane Experimental Technique, introducing the "Three Rs" — Replacement (preferring non-animal methods when they can achieve the same scientific aims), Reduction (obtaining comparable information from fewer animals, or more information from the same number), and Refinement (minimizing pain, suffering, or distress and enhancing welfare for animals that must be used) — as guiding principles for more ethical animal research. The Three Rs would go on to become explicit in animal-research legislation in many countries and remain foundational to the ethical design of experiments involving animals.[160] United Kingdom
1959 Concept Arthur Samuel coins the term "machine learning" while developing a self-improving computer program to play checkers, defining it as giving computers "the ability to learn without being explicitly programmed" — though this specific phrasing is now considered apocryphal and does not appear in Samuel's actual publications.[161] United States
1959–1961 Method Jack Kiefer and Jacob Wolfowitz develop formal optimal design theory, introducing criteria-based selection of experimental designs for maximum precision.[21] United States
1960 Industrial statistics Japanese engineer and statistician Genichi Taguchi publishes Design of Experiments for Engineers, advancing statistical methods for industrial experimentation. His work helps formalize what later would become known as Taguchi methods, integrating experimental design and statistical analysis to improve product quality, optimize manufacturing processes, and reduce costs in engineering and industry.[162] Japan
1960 Method George E. P. Box and Donald Behnken publish "Some New Three Level Designs for the Study of Quantitative Variables" in Technometrics, devising what becomes known as Box–Behnken designs — three-level experimental designs for response surface methodology that place each factor at one of three equally spaced values and are efficient enough to fit a quadratic model. The seven-factor design, whose estimation variance depends almost exactly on distance from the centre point, was found first; designs for other numbers of factors were subsequently derived to approximate this same rotatability property.[163][164] United States
1960 Experiment English psychologist Peter Wason conducts an experiment in which participants must discover a rule governing triples of numbers, given that (2,4,6) fits it. Participants overwhelmingly test only triples that confirm their current hypothesis rather than triples designed to falsify it, and rarely discover the true (very broad) rule, "any ascending sequence." Wason interprets the results as showing a preference for confirmation over falsification in hypothesis testing, coining the term "confirmation bias."[165] United Kingdom
1960 Method Donald Thistlethwaite and Donald T. Campbell introduce the regression discontinuity design in "Regression-Discontinuity Analysis: An alternative to the ex post facto experiment," published in the Journal of Educational Psychology, applying it to evaluate the effect of merit-based scholarship programs on students' career plans. The design estimates a local causal treatment effect by comparing outcomes for units lying just above and just below a predetermined threshold that determines treatment assignment (e.g. a scholarship awarded only to students scoring above a cutoff), exploiting the idea that units on either side of a narrow cutoff should otherwise be comparable even without formal random assignment. It becomes a widely used quasi-experimental design across economics, political science, epidemiology, and psychology.[166] United States
1960 Method The multiple baseline design is first reported in basic operant research, later applied in the late 1960s to human experimental subjects in response to practical and ethical concerns about withdrawing apparently successful treatments once established — a design in which treatment is introduced to different behaviors, individuals, or settings at staggered intervals, so that treatment-timed changes (rather than chance) become attributable to the intervention.[167] United States
1961 Concept George E. P. Box and J. S. Hunter introduce the concept of resolution in their paper "The 2^(k-p) Fractional Factorial Designs," providing a way to measure the degree to which a fractional factorial design avoids aliasing between main effects and important interactions — a design has resolution R if every effect involving p factors is unaliased with every effect having fewer than R−p factors. This concept becomes a standard way of summarizing and comparing the aliasing properties of fractional designs.[168] United States
1961 Experiment Ecologist Joseph H. Connell conducts field experiments on rocky intertidal shores manipulating the presence of competing barnacle species, becoming an influential and widely cited example of field experimentation in ecology and helping establish an experimental research tradition on rocky seashores that continued through the 1980s.[169] United States
1961 Concept The term nocebo (Latin nocēbō, "I shall harm", from noceō, "I harm")[170] is coined by Walter Kennedy to denote the counterpart to the use of placebo (Latin placēbō, "I shall please", from placeō, "I please"; a substance that may produce a beneficial, healthful, pleasant, or desirable effect). Kennedy emphasized that his use of the term "nocebo" refers strictly to a subject-centered response, a quality inherent in the patient rather than in the remedy".[171] United States
1962 Policy The Kefauver Harris Amendment requires proof of efficacy through controlled trials before drug approval, institutionalizing experimental design in regulation. United States
1962 Experiment Researchers including Walter N. Pahnke conduct the Marsh Chapel Experiment at Boston University, a double-blind study giving divinity students either the psychedelic substance psilocybin or an active placebo — a large dose of niacin chosen specifically because it produces noticeable physical sensations that could lead control subjects to believe they too had received a psychoactive drug. The design becomes an influential early example of using an active rather than inert placebo to counter unblinding in trials of drugs with strong subjective effects.[172] United States
1962 Method Japanese statistician Kazumasa Kôno publishes "Optimum designs for quadratic regression on k-cube" in the Memoirs of the Faculty of Science, Kyushu University, deriving optimal experimental designs for quadratic response surface methodology models that require fewer experimental runs than George E. P. Box's earlier central-composite designs. Together with later work by Jack Kiefer (mathematician), the Kôno–Kiefer analysis explains why such optimal designs, despite their mathematical construction as probability measures over continuous spaces, can be supported on a small number of discrete points closely resembling the traditional designs already used in practice.[173] Japan
1962 Experiment Vernon L. Smith publishes An Experimental Study of Competitive Market Behavior in the Journal of Political Economy. Using controlled laboratory market experiments with human participants, he tests hypotheses of neoclassical competitive theory and demonstrates price convergence toward equilibrium, helping establish experimental economics as a systematic empirical research methodology.[174] United States
1962 Method British statistician John Nelder proposes a set of systematic, circular experimental designs as an alternative to the replicated, full factorial spacing experiments. These designs, known as the Nelder 'wheel' design, are developed to address limitations related to space and plant material. The design consists of a circular plot with concentric circumferences radiating outward, connected by spokes that extend from the center to the farthest circumference. Trees are planted at the intersections of spokes and circumferences within the plot.[175] United Kingdom
1963 Concept Campbell and Stanley discuss design according to the categories of preexperimental designs, experimental designs, and quasi-experimental designs.[176] United States
1963 Method J.J. Boren introduces "repeated acquisition of new behavioral chains" as a single-subject research design in the American Psychologist, offering an alternative to reversal designs for studying behaviors that cannot practically or ethically be reversed once learned — instead repeatedly measuring how quickly a subject acquires a new behavioral sequence under different experimental conditions.[177] United States
1964 Policy The World Medical Association adopts the Declaration of Helsinki, further developing the informed-consent and human-experimentation principles first codified in the 1947 Nuremberg Code. The Declaration becomes the foundational reference document underlying the guidelines of research ethics committees worldwide.[178] Finland
1965 Concept Leslie Kish coins the term "design effect" in his book Survey Sampling, proposing the general definition as the ratio of the variance of an estimator under a given (often complex) sampling design to its variance under simple random sampling of the same size — along with formulas for the design effect under cluster sampling (incorporating intraclass correlation) and under unequal-probability sampling. These become known as "Kish's design effect" and remain foundational to survey methodology.[179] United States
1966 Concept Hanan Selvin and Alan Stuart describe "data-dredging procedures" in survey analysis, formalizing an early statistical critique of selecting which explanatory variables to retain in a model based on the same data used to test them. Using the metaphor of fish that "don't fall through the net," they argue that variables retained after such a selection process are systematically biased toward appearing more significant than they truly are, altering the validity of standard statistical tests applied afterward.[180] United Kingdom
1966 Publication American psychologist Robert Rosenthal (psychologist) publishes Experimenter Effects in Behavioral Research, a foundational study establishing that researchers can subtly and unconsciously communicate their expectations to experimental subjects, biasing outcomes toward those expectations — later demonstrated experimentally by showing that researchers told to expect positive ratings from subjects obtained significantly more positive data than researchers told to expect negative ratings, using the identical task.[181] United States
1971 Publication Raymond H. Myers publishes Response Surface Methodology, a textbook establishing widely used practical guidelines for central composite design parameters — including standard methods for selecting the axial-point distance α (orthogonal and rotatable design criteria) — that remain in use in statistical software decades later.[182] United States
1972 Publication Herman Chernoff writes an overview of optimal sequential designs[183] In the design of experiments, optimal designs is a class of experimental designs that are optimal with respect to some statistical criterion. United States
1973 Concept Belgian mathematician Philippe Delsarte's thesis, "An Algebraic Approach to the Association Schemes of Coding Theory," recognizes and develops the deep connections between association schemes and both coding theory and design theory, becoming widely regarded as the most important contribution to association scheme theory after its statistical origins in design of experiments.[184] Belgium
1973 Method Donald Rubin publishes "Matching to Remove Bias in Observational Studies" in Biometrics, formalizing statistical matching — pairing treated units in an observational study with non-treated units sharing similar observable characteristics — as a technique for reducing confounding bias in estimated treatment effects when random assignment is unavailable. The paper predates and lays groundwork for Rubin's later development of propensity score matching with Paul Rosenbaum in 1983.[185] United States
1974 Concept Donald Rubin formalizes the potential outcomes framework, defining causal effects using counterfactual comparisons between treated and untreated states for the same unit.[186] United States
1974 Method R.H.F. Denniston solves a problem posed by James Joseph Sylvester in 1860 as an extension of Kirkman's schoolgirl problem: whether thirteen disjoint Steiner triple systems S(2,3,15) exist, such that "Kirkman's schoolgirls" could march in triples for an entire 13-week term without any pair of girls being grouped together twice. Denniston's computer search, run for seven hours on an Elliott 4130 computer at the University of Leicester, finds a valid week-one solution from which all subsequent weeks are generated by a fixed relabeling scheme; the number of non-isomorphic solutions to Sylvester's problem remains unknown.[187] United Kingdom
1975 Policy The first revision to the Declaration of Helsinki ("Helsinki II") becomes the first international guideline to require that a research protocol for human experimentation be approved by an independent ethics committee before the study may proceed — formally establishing the ethics-committee review model that would become standard practice internationally.[188] Global
1975 Concept Dijen K. Ray-Chaudhuri and R. M. Wilson prove a generalization of Fisher's inequality to t-designs, showing that in a 2s-(v, k, λ) design the number of blocks is at least the binomial coefficient v choose s — extending the classical Fisher's inequality (bv for balanced incomplete block designs) to a broader class of combinatorial designs.[189] United States
1975 Concept A.C. Atkinson and V.V. Fedorov publish "The design of experiments for discriminating between two rival models" in Biometrika, introducing T-optimality — a criterion for constructing experimental designs specifically intended to maximize the statistical discrepancy between two competing candidate models at the chosen design points, rather than to minimize estimation variance within a single assumed model. The criterion becomes particularly important in biostatistics applications supporting pharmacokinetics and pharmacodynamics, building on earlier work by David R. Cox and Atkinson on model-discrimination experiments.[190] United Kingdom
1975 Method Stuart Pocock and Richard Simon introduce minimisation in Biometrics, an adaptive stratified sampling technique for balancing treatment groups across multiple prognostic factors in clinical trials. Unlike traditional blocking (statistics), which requires a separate randomisation list for each combination of stratification factors — a number of lists that grows exponentially as more factors are added — minimisation calculates, for each new patient, the imbalance that would result from allocating them to each treatment group, sums the imbalance across all factors, and assigns the patient to whichever group minimises the overall imbalance (optionally with a random element retained). The method comes to be described by some as maintaining better balance than blocked randomisation, particularly as the number of stratification factors grows.[191] United Kingdom
1976 Method James Scheirer, William S. Ray, and Nathan Hare publish "The Analysis of Ranked Data Derived from Completely Randomized Factorial Designs" in Biometrics, introducing the Scheirer–Ray–Hare test, a non-parametric extension of the Kruskal–Wallis one-way analysis of variance for examining whether a measured outcome is affected by two or more factors, without requiring the assumption of a normal data distribution. Though more conservative in statistical power than parametric multi-factorial ANOVA and later described than most comparable variance-analysis tests, the Scheirer–Ray–Hare test finds continued use, particularly in the biological sciences, as a non-parametric alternative when normality assumptions cannot be met.[192] United States
1976 Publication Douglas C. Montgomery publishes Design and Analysis of Experiments, a comprehensive textbook on the design and analysis of experiments. The book covers a wide range of topics, including principles of experimental design, different types of experimental designs, analysis of experimental data, and use of experimental design in a variety of fields, such as agriculture, industry, and medicine.[193] United States
1976 Method Olli Miettinen formalizes the conditions under which the odds ratio of exposure in a case–control study can be used to estimate relative risk, refining earlier work by Jerome Cornfield that had shown this approximation held specifically when the disease outcome under study is rare.[194] United States
1976 Method James Scheirer, William S. Ray, and Nathan Hare publish "The Analysis of Ranked Data Derived from Completely Randomized Factorial Designs" in Biometrics, introducing the Scheirer–Ray–Hare test, a non-parametric extension of the Kruskal–Wallis one-way analysis of variance for examining whether a measured outcome is affected by two or more factors, without requiring the assumption of a normal data distribution. Though more conservative in statistical power than parametric multi-factorial ANOVA and later described than most comparable variance-analysis tests, the Scheirer–Ray–Hare test finds continued use, particularly in the biological sciences, as a non-parametric alternative when normality assumptions cannot be met.[195] United States
1977 Method The concept of Pocock boundary is introduced by the medical statistician Stuart Pocock.[196] United Kingdom
1977 Concept Latvian engineer Vilnis Eglājs proposes a space-filling multifactor experimental design technique in Russian-language literature that is later recognized as equivalent to Latin hypercube sampling — predating Michael McKay's independently-derived and far more widely cited 1979 description of the same technique by two years, though Eglājs's work remains comparatively obscure in the Western statistical literature.[197] Latvia
1977 Method Stuart Pocock introduces group sequential methods for the design and analysis of clinical trials, in which patient entry is divided into equal-sized groups so that repeated significance tests on the accumulated data after each group determine whether the trial continues.[198] United Kingdom
1978 Concept Donald Rubin introduces the concept of "ignorable assignment mechanisms" in causal inference, formalizing the condition under which the way individuals were assigned to treatment groups can be disregarded during statistical analysis, given everything else recorded about them — a refinement of his earlier Rubin causal model potential-outcomes framework, and part of what becomes known as the Neyman-Rubin causal inference model.[199] United States
1978 Concept According to Box et al., experimental design refers to the systematic layout of combinations of variables. The layouts in the case of concepts are test concepts or test vignettes.[200] United States
1978 Concept Ulrich Krengel and Louis Sucheston (with David J. H. Garling) formulated the Prophet Inequality in optimal stopping theory. They show that a gambler observing sequential random rewards can secure at least half the expected payoff of a “prophet” who knows all outcomes in advance, establishing a foundational result in probability theory and decision processes.[201] United States
1979 Method Marvin Zelen publishes his new method, which would later be called Zelen's design.[202][203] United States
1979 Publication Thomas D. Cook and Donald T. Campbell publish Quasi-experimentation: Design & Analysis Issues for Field Settings, extending Campbell and Julian C. Stanley's 1963 categorization of experimental, quasi-experimental, and pre-experimental designs into a comprehensive treatment of quasi-experimental methodology for applied field research where randomization is impractical or unethical. The book becomes a foundational reference for quasi-experimental design across social science, public health, education, and policy analysis, and is substantially revised and expanded by Cook, Campbell, and William R. Shadish in 2002.[204] United States
1979 Method Peter C. O'Brien and Thomas R. Fleming publish "A Multiple Testing Procedure for Clinical Trials" in Biometrics, introducing the O'Brien–Fleming boundary for group sequential design clinical trial monitoring. The boundary sets a very conservative (high) significance threshold at early interim analyses, becoming progressively less stringent as the trial proceeds until it approaches the nominal significance level (e.g. 0.05) at the final analysis — protecting against premature stopping due to random early fluctuations while still preserving the overall Type I error rate specified for the trial. Despite requiring the number and timing of interim analyses to be prespecified, the O'Brien–Fleming boundary becomes one of the most widely used methods for monitoring clinical trials.[205] United States
1979 Method Michael McKay at Los Alamos National Laboratory makes a significant contribution to the field of statistical sampling by introducing the concept of latin hypercube sampling.[206] United States
1979 Concept John C. Gittins publishes "Bandit Processes and Dynamic Allocation Indices," proving that the optimal solution to the multi-armed bandit problem — first posed as an unsolved question in Herbert Robbins's 1952 paper on sequential experimental design — takes the form of an index policy: at each stage, choosing the option (or "arm") with the highest computable "dynamic allocation index," now known as the Gittins index. The result resolves optimal-stopping questions in clinical trial design that had been open since the 1940s, and independent economist Martin Weitzman would establish the equivalent result in economics the same year.[207] United Kingdom
1980 Concept Researchers publish long-term results of the Coronary Drug Project, a study of drugs for long-term treatment of coronary heart disease in men, reporting that participants in the placebo group who adhered to their placebo regimen as instructed showed nearly half the mortality rate of those who did not adhere — despite the placebo itself being pharmacologically inert. The finding becomes a widely cited illustration of the "healthy adherer" effect, in which apparent treatment benefits attributed to adherence may instead reflect underlying differences between compliant and non-compliant patients (health consciousness, psychological effects of following a protocol, or general diligence), a confound later replicated in women with nearly 2.5 times greater survival among placebo-adherent patients.[208] United States
1980 Concept Statistician John Tukey writes about the choice between confirmatory analysis (testing or rejecting existing hypotheses) and exploratory analysis (searching for new hypotheses) in statistical practice, examining how practicing statisticians decide between the two modes of reasoning at different stages of research — an influential early articulation, within statistics specifically, of the same exploratory/confirmatory distinction later applied more broadly to human reasoning by psychologists Jennifer Lerner and Philip Tetlock in 2002.[209] United States
1981 Concept Allen Neuringer first proposes the idea of using single case designs (sometimes referred to as n-of-1 trials) for self-experimentation.[210] United States
1982 Literature British statistician George Box publishes Improving Almost Anything: Ideas and Essays, which gives many examples of the benefits of factorial experiments.[211] United States
1983 Concept Donald Rubin and Paul R. Rosenbaum define the stronger condition of a treatment assignment being "strongly ignorable" in their paper on propensity score matching, establishing propensity scores — the probability of treatment assignment given observed covariates — as central to estimating causal effects from observational data when random assignment is unavailable.[212] United States
1983 Method K.K. Gordon Lan and David L. DeMets publish "Discrete Sequential Boundaries for Clinical Trials" in Biometrika, proposing a method that allows the boundary values of a group sequential trial to be allocated dynamically as the study progresses, addressing a key limitation of the O'Brien–Fleming boundary and other fixed methods that require prespecifying the number of interim analyses and the proportion of total information used at each one.[213] United States
1984 Method French pharmacologist Bernard Bégaud describes the challenge–dechallenge–rechallenge protocol as one of the standardized methods used in France for assessing adverse drug reactions, monitoring whether an adverse event resolves on withdrawal of a medication (dechallenge) and recurs on its re-administration (rechallenge) — a design suited to idiosyncratic, individual-level reactions where conventional population-level statistical testing is unsuitable.[214] France
1984 Concept Stuart Hurlbert publishes a paper in Ecological Monographs where he analyzes 176 experimental studies in ecology. He discovers that 27% of these studies suffer from 'pseudoreplication,' meaning they use statistical testing in situations where treatments are not replicated or replicates were not independent. When considering only studies that use inferential statistics, the percentage of pseudoreplication increases to 48%. To address this issue, Hurlbert suggests interspersing treatments in experiments, even if it means sacrificing randomized samples, particularly in smaller experiments. This approach aims to overcome the problem of pseudoreplication in ecological studies.[215] United States
1986 Experiment Robert LaLonde finds that findings of econometric procedures assessing the effect of an employment program on trainee earnings do not recover the experimental findings. This is considered to be the start of experimental benchmarking in social science.[216] United States
1986 Concept Fred N. Kerlinger describes the MAXMINCON principle, emphasizing maximizing systematic variance, controlling extraneous variance, and minimizing error variance in experimental design.[176] United States
1987 Publication Australian mathematician Anne Penfold Street publishes Combinatorics of Experimental Design, a textbook on combinatorial methods in experimental design.[217] Australia
1987 Concept S.K. Wang and Anastasios A. Tsiatis publish "Approximately optimal one-parameter boundaries for group sequential trials" in Biometrics, introducing a general parameterized family of stopping boundaries of which both the O'Brien–Fleming boundary (1979) and the Pocock boundary (1977) are shown to be special cases.[218] United States
1987 Publication Australian mathematician Anne Penfold Street publishes Combinatorics of Experimental Design, a textbook on combinatorial methods in experimental design.[219] Australia
1987 Publication Australian mathematician Anne Penfold Street and her daughter, statistician Deborah Street, publish Combinatorics of Experimental Design, a textbook connecting the combinatorial mathematics of block designs, Latin squares, and factorial designs to their applications in statistics. Reviewers praised the book for making the combinatorial side of experimental design accessible to statisticians, with Marshall Hall calling it "very readable" and "very satisfying," though some noted it omitted certain topics covered by more comprehensive contemporary texts.[220][221] Australia
1988 Publication Roger Mead publishes The Design of Experiments: Statistical Principles for Practical Applications, a textbook presenting practical principles of experimental design and analysis.[222] United Kingdom
1988 Method Organizational psychologists Miriam Erez and Gary P. Latham, in a dispute over the effect of participation on goal commitment and performance in goal setting research, design four experiments together with Edwin Locke serving as a neutral third party to resolve their disagreement — one of the earliest documented modern examples of what would later be termed adversarial collaboration, though the term itself was not yet coined.[223] United States
1988 Organization Stat-Ease releases its first version of Design–Expert, a statistical software package specifically dedicated to performing design of experiments (DOE), offering comparative tests, screening, characterization, and optimization tools.[224] United States
1989 Publication Perry D. Haaland publishes Experimental Design in Biotechnology, presenting statistical experimental design and analysis as a problem-solving tool in biotechnology.[225][226][227] United States
1989 Method Jerome Sacks, William Welch, Toby Mitchell, and Henry Wynn publish a landmark paper summarizing the Bayesian statistical framework for computer experiments, modeling a deterministic computer simulation's output as an unknown function of its inputs and placing a Gaussian process prior over that function — establishing an approach to designing and analyzing simulation experiments distinct from classical experimental design for physical systems, where criteria like replication and A/D-optimality (suited to parametric models with random error) do not directly apply.[228] United States
1989 Publication Jerome Sacks and collaborators discuss statistical issues in the design and analysis of computer and simulation experiments, helping establish the field of computer experiments.[229] United States
1989 Concept C. W. H. Lam and collaborators complete a large-scale computer search establishing that no finite projective plane of order 10 exists — a case the Bruck–Ryser–Chowla theorem's necessary conditions do not themselves rule out, showing that theorem's criteria, while powerful, are not sufficient to determine existence in general. The proof, which combined coding theory with an extensive computer search, was notable enough to draw coverage in The New York Times questioning whether a computer-assisted proof that no human could fully verify by hand should count as a mathematical proof.[230][231] Canada
1990 Organization Regulatory authorities and pharmaceutical industry representatives from Europe, Japan, and the United States establish the International Conference on Harmonisation of Technical Requirements for Registration of Pharmaceuticals for Human Use (ICH), a joint initiative to harmonize clinical trial protocol standards across jurisdictions — aiming to ensure quality, safety, and efficacy in drug development while preventing unnecessary duplication of human trials and minimizing animal testing.[232] Global
1991 Organization The first International Data Farming Workshop takes place, part of a series later organized by the SEED Center for Data Farming at the Naval Postgraduate School; over the following decades, more than 16 additional workshops would be held, with participation from countries including Canada, Singapore, Mexico, Turkey, and the United States, each assigning teams of researchers to apply data farming techniques to specific domains such as robotics, homeland security, and disaster relief.[233][234] Global
1993 Concept Judea Pearl introduces the Back-Door criterion, a graphical condition using causal graphical model for identifying a sufficient set of variables to adjust for in order to obtain an unbiased estimate of a causal effect — providing a formal, graph-based alternative to the counterfactual definitions of confounding developed in epidemiology, later shown to be formally equivalent to them.[235] United States
1993 Policy The Council for International Organizations of Medical Sciences (CIOMS), a body established by the World Health Organization, first publishes International Ethical Guidelines for Biomedical Research Involving Human Subjects, formally requiring ethics committees and focusing particular attention on research practice in developing countries. Though the guidelines carry no legal force, they become influential in shaping national regulations governing ethics committees.[236] Global
1994 Method American psychiatrist Peter Breggin applies the challenge–dechallenge–rechallenge (CDR) protocol — administering, withdrawing, then re-administering a medication while monitoring for adverse effects — to investigate a suspected association between fluoxetine (Prozac) and suicidal ideation. Breggin had observed that only certain individuals responded to the medication with increased suicidal thoughts; given the low occurrence rate of this reaction, conventional statistical testing across a study population was considered inappropriate, making the individual-level CDR protocol a more suitable design. Eli Lilly and Company subsequently adopted the CDR protocol, rather than a randomized controlled trial, when testing for increased suicide risk associated with the drug.[237] United States
1994 Experiment David Card and Alan Krueger publish a study using the difference in differences design to evaluate the employment effects of New Jersey's April 1992 minimum wage increase, comparing fast-food employment in New Jersey against neighboring Pennsylvania (used as a control not subject to the wage increase) before and after the change. Contrary to standard economic theory's prediction that a minimum wage increase would reduce employment, they find no such decrease — an early, influential demonstration of using a natural experiment and a control region to estimate a causal treatment effect from observational data. Card received the 2021 Nobel Memorial Prize in Economic Sciences in part for this and related work.[238] United States
1994 Method The Neyer-d optimal test is first described by Barry T. Neyer.[239] United States
1995 Method Chris Nachtsheim and Ruth Meyer introduce the coordinate exchange algorithm, enabling computational generation of optimal experimental designs.[24] United States
1995 Concept Leslie Kish proposes the "Design Effect Factor" (Deft), a refinement of his 1965 design effect that uses simple random sampling with replacement in the denominator rather than without, arguing this better isolates the effect of sampling design from the nuisance of finite-population correction and is simpler to use directly in confidence-interval calculations.[240] United States
1995 Publication Kathryn Chaloner and Isabella Verdinelli publish "Bayesian Experimental Design: A Review" in Statistical Science, a widely cited survey establishing the now-standard approach of assuming approximate normality of posterior probabilities in order to calculate expected utility using linear theory — becoming a standard reference point for later work in Bayesian experimental design.[241] United States
1996 Ethical oversight International Conference on Harmonisation (ICH) guidelines for Good Clinical Practice (GCP) established.[242] Global
1996 Method Colombian-Canadian physician Alex Jadad, working as a Research Fellow at Oxford's Pain Relief Unit, and colleagues introduce the Jadad scale (also known as Jadad scoring or the Oxford quality scoring system) in an appendix to a paper on blinding in randomized clinical trials. The instrument scores a trial report from zero to five based on three yes/no questions covering randomization, double-blinding, and reporting of withdrawals and dropouts. Despite later criticism for over-emphasizing blinding, showing low inter-rater consistency, and omitting allocation concealment, the scale becomes the most widely used trial-quality assessment tool worldwide, with its seminal paper cited in over 25,000 scientific works by 2024.[243] United Kingdom
1996 Method Stat-Ease releases Design–Expert Version 5, the first version of the software designed for Microsoft Windows, broadening accessibility of dedicated design-of-experiments software beyond earlier DOS-based tools.[244] United States
1996 Policy The Consolidated Standards of Reporting Trials (CONSORT) Statement is first published, the product of a 1995 Chicago meeting merging two independent efforts to improve randomized-trial reporting: the 1993 Ottawa meeting's Standardized Reporting of Trials (SORT) proposal and the concurrent Asilomar Working Group's recommendations from California. Convened at the suggestion of JAMA's Drummond Rennie, the merged CONSORT Statement provides a standardized checklist and participant flow diagram intended to reduce bias and aid critical appraisal of trial reports.[245] United States
1997 Concept Computer scientist Tom M. Mitchell proposes a widely cited formal definition of machine learning: "A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P if its performance at tasks in T, as measured by P, improves with experience E." — offering an operational, experimentally-grounded definition rather than an aspirational one, directly framing learning in terms of measurable performance under controlled tasks.[246] United States
1998 Policy The U.S. Department of Health and Human Services announces a policy change to the population standard used for age-adjusting death rates in its publications.[247] United States
1998 Concept A. Brandstein and G. Horne coin the term "data farming" in conjunction with the U.S. Marine Corps' Project Albert, describing the use of large-scale designed computational experiments — running small agent-based simulation models thousands or millions of times at high-performance computing facilities — to "grow" data that can then be analyzed for insight into complex systems, distinguishing the approach from data mining (which extracts patterns from data one has no control over generating).[248] United States
1999 Method Rajeev Dehejia and Sadek Wahba re-examine Robert LaLonde's original 1986 dataset using additional non-experimental methods, arguing that when there is sufficient overlap between treated and untreated subject pools and unobservable covariates do not substantially impact outcomes, non-experimental methods can in fact estimate treatment effects accurately — offering a more optimistic counterpoint to LaLonde's original benchmarking critique.[249] United States
1999 Concept Basili et al use the term family of experiments to refer to a group of experiments that pursue the same goal and whose results can be combined into joint—and potentially more mature—findings than those that can be achieved in isolated experiments.[250] United States
2000 (January 19) Publication A First Course in Design and Analysis of Experiments.[251] United States
2001 Method The World Health Organization publishes a new standard population for age standardization, intended to allow health statistics from different countries to be compared despite differing population age profiles.[252] Global
Late 1990s Organization Bradley Jones joins JMP (software), contributing to the development of advanced DOE software tools for engineers and researchers.[24] United States
Late 20th century Concept Fisher’s principles of randomization, replication, and blocking become standard features of statistically rigorous experiments in the biological and biomedical sciences.[19] Global
Late 20th century Concept Recognition that all experiments are inherently designed, with emphasis on the importance of proper planning to avoid wasted resources and invalid results.[21] Global
Late 20th century Method Advances in computing and algorithms enable practical implementation of optimal design, expanding its use in scientific and industrial applications.[21] Global
21st century Application Designed experiments are widely applied across service sectors including finance, business operations, and government, reflecting the broad adoption of statistical experimentation.[22] Global
2001 Method Daniel Kahneman independently develops a protocol for adversarial collaboration roughly a decade after Erez, Latham, and Locke's earlier unnamed example, and may have been the first to use the term itself.[253][254] United States
2001 Policy Following further CONSORT Group meetings in 1999 and 2000, a revised CONSORT Statement is published, updating the original 1996 recommendations for reporting parallel-group randomized trials in light of growing empirical evidence on reporting quality.[255] Global
2002 Policy The World Medical Association issues a "Note of Clarification" on Paragraph 29 of the Declaration of Helsinki, addressing controversy over the ethics of placebo-controlled trials. The note reaffirms that placebo-controlled methodology should generally be used only in the absence of an existing proven therapy, but carves out two exceptions where it may remain ethically acceptable even when proven therapy exists: when compelling and scientifically sound methodological reasons make a placebo necessary to determine efficacy or safety, or when the condition under investigation is minor and placebo recipients face no additional risk of serious or irreversible harm. The clarification becomes a widely cited reference point in ongoing debates over the ethics of withholding proven treatment from trial participants.[256] Global
2002 Method Jochen Musch and Karl Christoph Klauer, and separately Ulf-Dietrich Reips, describe the "seriousness check" for web-based psychological experiments: asking respondents at the start of an online study whether they intend to participate seriously or merely browse the pages, in order to identify and exclude low-quality data entries before analysis. Later research finds the technique to be a strong predictor of dropout — roughly 75% of respondents who say they only want to look at the pages subsequently drop out, versus only 10–15% of those who say they intend to seriously participate — and that overall 30–50% of visitors to a typical online study fail the check.[257][258] Germany
2002 Method Howard Bloom, Charles Michalopoulos, Carolyn Hill, and Ying Lei conduct a large-scale experimental benchmarking study of mandatory welfare-to-work programs, testing which non-experimental methods come closest to recovering experimentally estimated program effects. They conclude that none of the non-experimental methods tested approach the accuracy of a randomized experiment for recovering the parameter of interest, reinforcing concerns first raised by Robert LaLonde's 1986 benchmarking study.[259] United States
2002 Concept The terms exploratory thought and confirmatory thought are introduced by social psychologist Jennifer Lerner and psychology professor Philip Tetlock in their book Emerging Perspectives in Judgment and Decision Making.[260] United States
2003 Organization The Abdul Latif Jameel Poverty Action Lab is founded to scale randomized evaluations in development economics.[261] United States
2003 Method Japanese researcher S. Hirata designs the loose-string task, which becomes the standard apparatus for cooperative pulling experiments: a single string threaded through loops on a movable platform such that if only one participant pulls, the string comes loose and the reward becomes unretrievable, requiring genuinely coordinated pulling for success. This design, simpler and more broadly applicable across species than Crawford's original box-and-rope apparatus, is subsequently adopted in cooperative pulling studies of rooks, ravens, wolves, elephants, capuchins, and many other species.[262] Japan
2004 Policy The U.S. Food and Drug Administration introduces the Critical Path Initiative, aimed at addressing high attrition rates in the clinical phase of drug development and offering investigators more flexibility to identify optimal clinical benefit without compromising study validity; adaptive designs for clinical trials initially emerge under this regulatory framework.[263] United States
2004 Policy The CONSORT group publishes an extension to the CONSORT statement specifically for cluster randomised trials, addressing the additional reporting requirements these designs need — such as accounting for intraclass correlation — beyond standard individually randomised trial reporting guidelines.[264] United Kingdom
2005 Experiment Study determines that most clinical trials have unclear allocation concealment in their protocols, in their publications, or both.[265] United Kingdom
2005 Publication Stuart Pocock publishes the editorial "When (not) to stop a clinical trial for benefit" in JAMA, addressing the range of practical and ethical considerations bearing on the decision to halt a trial early when interim results favor the treatment group — beyond the purely statistical threshold his own Pocock boundary had provided in 1977 — including the risk that early, promising results may not persist or generalize, and the tension between statistical stopping rules and the ethical pressure to give patients access to an apparently superior treatment.[266] United Kingdom
2005 Concept John Ioannidis publishes "Why Most Published Research Findings Are False" in PLOS Medicine, arguing on theoretical and probabilistic grounds that a majority of published research claims across many fields are likely to be false, due to factors including small sample sizes, small effect sizes, flexibility in study design and analysis, and bias from financial or other interests. The paper becomes one of the most-cited and most-discussed articles in the history of the reproducibility crisis debate, predating and helping motivate the preregistration and open-science reforms of the following decade.[267] United States
2005 Publication Jack Kleijnen, Susan Sanchez, Thomas Lucas, and Thomas Cioppa publish "A User's Guide to the Brave New World of Designing Simulation Experiments" in INFORMS Journal on Computing, addressing a key limitation of early data farming work: initial reliance on brute-force full factorial designs meant only a small number of factors could be investigated due to the curse of dimensionality. The paper helps establish improved experimental designs specifically suited to large-scale simulation experiments, developed in collaboration between Project Albert and the Naval Postgraduate School's SEED Center for Data Farming.[268] United States
2006 Publication American psychologist Seth Roberts publishes The Shangri-La Diet, a popular diet book based on conclusions he drew from self-experimentation — informally applying the N-of-1 trial logic Allen Neuringer had proposed for self-experimentation in 1981. Roberts, who documented his self-experiments on his blog, becomes a prominent early figure associated with the later quantified self movement, in which the growing ease of personal data collection and analysis drives a proliferation of N-of-1-style personal experiments.[269] United States
2006 Policy The U.S. Food and Drug Administration issues its Guidance on Exploratory Investigational New Drug (IND) Studies, formally introducing "Phase 0" trials — optional, exploratory human microdosing studies that administer a single subtherapeutic dose of a candidate drug or imaging agent to a small number of subjects (10–15) to gather early pharmacokinetic data. Because doses are too low to produce any therapeutic effect, Phase 0 trials yield no safety or efficacy data, but let developers rank candidate drugs and make go/no-go decisions based on relevant human data rather than solely on animal models, before committing to a full Phase I trial. The designation, initially unusual, becomes generally adopted as standard practice in drug development.[270] United States
2006 Method Alicia Melis, Brian Hare, and Michael Tomasello design a cooperative pulling experiment that explicitly controls for social tolerance between partners — a variable earlier studies had not accounted for — by comparing captive chimpanzee pairs known to share food readily against pairs less inclined to do so. They find food-sharing tolerance strongly predicts cooperative success, resolving much of the inconsistency across earlier chimpanzee cooperative pulling studies and establishing social tolerance as a key confounding factor that subsequent cooperative pulling experiments across many species would need to control for.[271] Germany
2007 (1 April) Organization The National Research Ethics Service (NRES) launches in the United Kingdom, a body requiring principal investigators to obtain approval for proposed research studies involving human participants before proceeding — with unapproved studies prohibited. NRES describes its purpose as reviewing research proposals "to protect the rights and safety of research participants and enable ethical research which is of potential benefit to science and society," extending the institutional ethics-committee review model established internationally by the Declaration of Helsinki decades earlier into a dedicated national service. The word "National" is later dropped from the name at an unrecorded point, and its functions are absorbed into the NHS Health Research Authority's Research Ethics Service.[272] United Kingdom
2008 Method Economist Justin McCrary proposes the "density test" for regression discontinuity designs, examining whether the density of observations of the assignment variable is continuous at the treatment cutoff. A discontinuity in this density — for example, an unusually large number of students who "just barely" pass an exam relative to those who "just barely" fail — suggests that some participants may have been able to manipulate their treatment status, undermining the "as good as random" assumption the design's validity depends on. The test becomes a standard diagnostic check in applied regression discontinuity research.[273] United States
2008 Concept A meta-epidemiological study of 146 meta-analyses finds that randomized controlled trials with inadequate or unclear allocation concealment tend to show results biased toward beneficial treatment effects — but only when trial outcomes are subjective rather than objective, refining earlier, less qualified claims about the relationship between allocation concealment and bias.[274] United Kingdom
2009 Method Adversarial collaboration is recommended by Daniel Kahneman[275] and others as a way of resolving contentious issues in fringe science, such as the existence or nonexistence of extrasensory perception.[276] United States
2010 Concept Asbjørn Hróbjartsson and Peter C. Gøtzsche argue in a meta-analysis that observed placebo effects can arise from bias due to lack of blinding, emphasizing the importance of proper control and blinding in experimental design.[277] Denmark
2010 Policy The National Centre for the Replacement, Refinement and Reduction of Animals in Research (NC3Rs) publishes the ARRIVE guidelines (Animals in Research: Reporting In Vivo Experiments), a 20-item checklist for improving experimental design and reporting standards in animal research, modeled on the CONSORT statement for reporting randomized trials. The guidelines follow a 2009 NC3Rs review finding that most biomedical journals provided little guidance on animal research design and reporting, and that large proportions of published animal studies failed to state a hypothesis, describe animal characteristics, use randomization or blinding, or fully report statistical methodology.[278][279] United Kingdom
2010 Policy The U.S. Food and Drug Administration issues draft guidance on adaptive design for clinical trials, an early regulatory step preceding the agency's more comprehensive 2019 guidance.[280] United States
2010 Method Researchers at the European Bioinformatics Institute (EMBL-EBI), led by J. Malone, publish a paper describing the Experimental Factor Ontology (EFO), an open-access ontology for systematically modeling experimental and sample variables — covering disease, anatomy, cell type, cell lines, chemical compounds, and assay information — originally developed to describe experimental variables in EBI's Expression Atlas resource. EFO is built to interoperate with existing biomedical ontologies such as ChEBI and the Ontology for Biomedical Investigations, and becomes a cross-cutting resource used for curation, querying, and data integration across resources including Ensembl and ChEMBL.[281] United Kingdom
2010 Policy The CONSORT 2010 Statement is published, consisting of a 25-item checklist and participant flow diagram, alongside an "Explanation and Elaboration" document detailing the reasoning behind each recommendation. By this point the CONSORT Statement has been endorsed by over 600 journals and editorial groups, including The Lancet, BMJ, JAMA, and the New England Journal of Medicine, and has directly inspired parallel reporting-guideline initiatives for other study types, including STROBE for observational studies and PRISMA for systematic reviews.[282] Global
2011 Method Stefano M. Iacus, Gary King, and Giuseppe Porro introduce Coarsened Exact Matching (CEM) in the Journal of the American Statistical Association, a matching method they present as simpler, more statistically powerful, and monotonic imbalance bounding compared to earlier techniques such as propensity score matching.[283] United States
2010 Experiment The I-SPY 2 trial launches as an adaptive Phase 2 clinical trial platform for breast cancer, linking academic cancer centers, the FDA, the NCI, and pharmaceutical partners. Building on predictive biomarkers developed in its predecessor I-SPY 1 (2002–2006), the trial evaluates multiple experimental drug combinations against standard chemotherapy simultaneously, dropping ineffective regimens early and advancing promising ones to confirmatory trials — a widely cited real-world demonstration of adaptive experimental design accelerating drug development.[284] United States
2013 Method Alex Deng, Ya Xu, Ron Kohavi, and Toby Walker introduce CUPED (Controlled-experiment Using Pre-Experiment Data) at the WSDM international conference, a technique for improving the statistical sensitivity of online A/B tests by using data from before an experiment began to reduce the variance of outcome measurements — allowing web companies to detect smaller effects with the same sample size, or the same effects with less traffic. The technique becomes widely adopted in large-scale online experimentation at companies running many concurrent experiments on millions of users.[285] United States
2014 Concept Uri Simonsohn, Leif Nelson, and Joseph Simmons — the researchers behind the blog Data Colada — coin the term "p-hacking" to describe the practice of running many statistical analyses on the same data set and reporting only those that yield a statistically significant result, dramatically understating the true risk of false positives. The paper also introduces the "p-curve" technique for detecting p-hacking in published literature by examining the distribution of significant p-values across studies.[286] United States
2014 Concept Peter Keevash proves the existence of nontrivial Steiner systems for all t ≥ 6, resolving a long-standing open problem in design theory, in "The existence of designs." His proof is non-constructive, and as of 2019 no explicit Steiner systems are actually known for large values of t despite the existence proof.[287] United Kingdom
2015 Experiment Journalist John Bohannon deliberately conducts and publishes a fraudulent study claiming chocolate consumption accelerates weight loss, using p-hacking techniques — considering 18 different variables during testing until one produced a statistically significant result — to demonstrate publicly how easily scientific-sounding claims can be manufactured from real but hollow statistical analysis. The hoax was picked up uncritically by numerous media outlets before Bohannon revealed it was an intentional social experiment exposing weaknesses in science journalism and nutrition research.[288] Germany
2015 Concept Megan L. Head, Luke Holman, Rob Lanfear, Andrew T. Kahn, and Michael D. Jennions publish "The Extent and Consequences of P-Hacking in Science" in PLOS Biology, analyzing the distribution of p-values reported across a large sample of published papers to empirically estimate how widespread p-hacking actually is in scientific practice, rather than treating it as a purely theoretical risk. The study also examines factors — including pressure for early stopping mandated by some animal ethics boards when interim results reach significance — that can inadvertently encourage the practice.[289] Australia
2016 Concept A study finds that symphony orchestra auditions conducted behind a curtain, blinding judges to a performer's gender, increase the hiring of women — a widely cited real-world example of blinding used outside clinical or laboratory settings.[290] United States
2016 Method Susan Athey and Guido Imbens publish "Recursive partitioning for heterogeneous causal effects" in PNAS, pioneering machine learning techniques — adapting decision-tree methods to causal inference — for detecting and characterizing how treatment effects vary across subpopulations, addressing the external validity question of whether a treatment's effect generalizes across different subsets of people, times, and contexts rather than assuming a single homogeneous effect. Stefan Wager and Athey extend the approach using random forests in a 2018 follow-up paper.[291] United States
2017 Policy Nature Human Behaviour adopts the registered report publishing format, in which study proposals — including hypotheses and methods — are peer-reviewed and provisionally accepted for publication before data collection begins, with publication guaranteed regardless of outcome. The format is explicitly framed as a countermeasure to data dredging and HARKing (Hypothesizing After the Results are Known), shifting editorial evaluation from the results of research to the soundness of its questions and methods.[292] United Kingdom
2017 Organization The International Collaborative Network for N-of-1 Trials and Single-Case Designs (ICN) is established, co-chaired by Jane Nikles and Suzanne McDonald, as a global network of clinicians, researchers, and consumers interested in N-of-1 trials and single-case experimental designs; it grows to over 400 members across more than 30 countries.[293] Global
2017 Method Peter M. Aronow and Cyrus Samii publish "Estimating average causal effects under general interference, with application to a social network experiment" in the Annals of Applied Statistics, developing a general method for computing exposure probability matrices in experiments where SUTVA (the stable unit treatment value assumption) does not hold — such as experiments conducted over social networks where treated units can influence untreated neighbors — enabling inverse probability weighting approaches, including the Horvitz–Thompson estimator, to correct for the resulting bias when estimating average treatment and spillover effects.[294] United States
2017 Concept A historical case study documents the subversion of allocation concealment in a randomized controlled trial, part of a broader body of evidence that sealed-envelope and even centralized allocation concealment methods remain vulnerable to circumvention by study personnel — including researchers opening envelopes prematurely, holding them up to light, or keeping lists of previous allocations (reported by up to 15% of study personnel surveyed).[295] United Kingdom
2018 Experiment Brett Gordon, Florian Zettelmeyer, Neha Bhargava, and Dan Chapsky use data from large-scale field experiments on Facebook advertising to test whether standard observational methods — including propensity score matching, stratification, and regression adjustment — can recover the true causal effects of online ads on checkout, registration, and page-view outcomes. Despite the unusually rich variation available in social-media advertising data, they find observational methods are unable to accurately recover the causal effects established by the underlying randomized experiments, providing a striking contemporary demonstration of the same experimental-benchmarking concerns first raised by LaLonde in 1986.[296] United States
2018 Concept A study by Brian Nosek and colleagues proposes preregistration as a safeguard against p-hacking: researchers submit their data-analysis plan to a journal before beginning data collection, precluding after-the-fact manipulation of the analysis to reach statistical significance.[297] United States
2018 Concept Kenneth Schulz and colleagues publish a retrospective account of the introduction and adoption of the term "allocation concealment," distinguishing it from the related but distinct concept of blinded experiment: allocation concealment prevents foreknowledge of treatment assignment before randomization (addressing selection bias), while blinding conceals group identity after allocation (addressing ascertainment bias). The authors argue the earlier common term "randomization blinding" had confusingly conflated these two distinct sources of bias.[298] United States
2019 Recognition Nobel Prize in Economics awarded to Abhijit Banerjee, Esther Duflo, and Michael Kremer for experimental approaches to alleviating global poverty Sweden [299]
2019 Policy The U.S. Food and Drug Administration provides guidance on the use of adaptive designs in clinical trials, formalizing regulatory standards for flexible experimental designs.[300] United States
2019 Concept Gary King and Richard Nielsen publish "Why Propensity Scores Should Not Be Used for Matching" in Political Analysis, arguing that propensity score matching — developed by Rubin and Rosenbaum in 1983 and widely adopted since — actually increases model dependence, bias, and inefficiency relative to other matching methods, and should no longer be the default recommendation. The paper marks a significant reversal in applied best practice among political scientists, economists, and other users of observational causal-inference methods.[301] United States
2019 Concept Aleksander Fabijan, Jayant Gupchup, Somit Gupta, Jeff Omhover, Wen Qin, Lukas Vermeer, and Pavel Dmitriev publish "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments" at the ACM SIGKDD conference, formalizing sample ratio mismatch (SRM) — a statistically significant discrepancy between the expected and actual ratio of treatment and control group sizes in an experiment, often caused by failures in randomization or measurement instrumentation in online A/B testing — and proposing a taxonomy and detection rules of thumb, typically using a chi-squared goodness-of-fit test, to help practitioners identify SRM and avoid drawing conclusions from biased experimental data.[302] United States
2020 Experiment In response to the COVID-19 pandemic, the World Health Organization launches the Solidarity trial, European researchers launch the Discovery trial, and UK researchers launch the RECOVERY Trial — each a large-scale, multi-arm adaptive design clinical trial of candidate treatments for hospitalized patients with severe COVID-19 infection. The adaptive designs allow ineffective experimental treatments to be dropped quickly and replaced with others as evidence accumulates, letting researchers adjust trial parameters in near real time rather than waiting for a trial's predetermined endpoint, and become widely cited examples of adaptive trial methodology deployed at unprecedented speed and scale during a public health emergency.[303][304] Global
2020 Policy An international working group supported by the National Centre for the Replacement, Refinement and Reduction of Animals in Research publishes ARRIVE 2.0, a revision of the 2010 ARRIVE guidelines splitting the original 20-item checklist (effectively 38 items counting sub-items) into an "Essential 10" checklist of basic minimum requirements and a "Recommended Set" of 11 additional items, in response to studies finding the original guidelines — despite endorsement by over 600 journals by 2016 and over 1,000 by 2020 — had been largely ignored by researchers and made little measurable impact on reporting quality.[305][306] United Kingdom
2020 Method Researchers Elias Garcia-Pelegrin, Alexandra Schnell, Clive Wilkins, and Nicola Clayton argue in Science that magic tricks offer a novel approach to hypothesis testing and experimental design in the study of animal cognition: presenting a trick that reliably fools humans to a nonhuman animal, then using its behavioral response (e.g. prolonged looking time, taken as an indicator of surprise) to probe which cognitive blind spots — such as expectations about object permanence or attention control — the animal shares with humans. The approach faces practical challenges distinct from conventional experimental designs, including getting animals to attend to a human demonstrator and inferring "surprise" without verbal report.[307] United Kingdom
2020 Concept Jérôme Adda, Christian Decker, and Marco Ottaviani publish "P-hacking in clinical trials and how incentives shape the distribution of results across phases" in PNAS, providing empirical evidence that the distribution of statistically significant results in clinical trials shifts in ways consistent with p-hacking as financial and career incentives change across different phases of drug development — connecting the general phenomenon of p-hacking directly to the economics of pharmaceutical research.[308] Italy
2021 Experiment Garcia-Pelegrin, Schnell, Wilkins, and Clayton perform three sleight-of-hand magic tricks — palming, the French drop, and fast pass — on six Eurasian jays, an early empirical test of the magic-based experimental paradigm they had proposed the previous year. The jays are not deceived by palming or the French drop, both of which depend on human hand movements setting expectations about object location, but are successfully deceived by the faster "fast pass" technique — the first demonstrated case of a magic trick fooling a nonhuman animal.[309] United Kingdom
2020 Concept Gonville and Caius College, Cambridge removes a stained-glass window honoring Ronald Fisher — which had depicted a 7×7 Latin square in reference to his Design of Experiments — because of Fisher's association with eugenics. The removal, following a vote by the college's governing body, becomes a widely covered episode in the broader reckoning of statistics and genetics with the eugenic commitments of several of the field's founding figures.[310] United Kingdom
2021 Method Researchers Marianne Gunderson, Kristian A. Bjørkelo, and Jill Walker Rettberg, together with a team of larp designers led by Anita Myhre Andersen, stage Sivilisasjonens venterom ("Civilization's Waiting Room") in Bergen, Norway — a live-action roleplaying game (larp) designed as a research methodology in its own right, intended to let participants practice ethical decision-making around emerging surveillance and machine-vision technologies (facial recognition, deepfakes, VR) within a fictional AI-governed society. The project is analyzed in subsequent scholarship as a "mimetic method" related to design fiction, illustrating how immersive, participatory fiction can function as a qualitative research design distinct from conventional controlled experimentation.[311] Norway
2022 Organization Philip Tetlock and Cory Clark propose adversarial collaboration as a vehicle for scientific self-correction, arguing it helps expose false claims by forcing exploration of rival hypotheses; their work leads to the University of Pennsylvania School of Arts & Sciences establishing the Adversarial Collaboration Project to formally support and encourage the approach across research questions.[312][313] United States
2025 Policy The CONSORT Group issues the CONSORT 2025 checklist, superseding the 2010 statement by adding seven new items, revising three, removing one, and introducing a new open-science reporting section covering trial registration, protocol and statistical-analysis-plan access, data sharing, and conflict-of-interest disclosures.[314] Global
2025 Concept Physician-scientist Pavlos Msaouel publishes "The curious rise of randomised non-comparative trials" in Significance, critiquing a design — traced by a companion meta-epidemiological review to oncology studies dating back to 2002 — in which participants are randomized to different treatment arms, but each arm is compared only against a historical control or predefined benchmark rather than against each other, functioning in practice as multiple concurrent single-arm studies. Msaouel argues the term "randomisation" in these designs is largely "talismanic," misleadingly suggesting the methodological benefits of a true randomized controlled trial despite providing no formal between-arm comparison; a companion analysis by Alexander D. Sherry, Msaouel, and Ethan B. Ludmir finds that despite this, roughly half of published randomised non-comparative trials still report some form of comparison between arms anyway.[315][316] United States
2020s Method Machine learning is integrated with experimentation to estimate heterogeneous treatment effects and optimize interventions Global [317]

Numerical and visual data

Google Scholar

The following table summarizes per-year mentions on Google Scholar as of December 14, 2021.

Year "experimental design"
1900 30
1910 17
1920 13
1930 19
1940 62
1950 425
1960 1,590
1970 6,240
1980 11,400
1990 17,000
2000 53,200
2010 162,000
2020 90,600

The chart below shows Google Trends data for Design of experiments (Topic), from January 2004 to December 2021, when the screenshot was taken. Interest is also ranked by country and displayed on world map.[318]

Google Ngram Viewer

The chart below shows Google Ngram Viewer data for Design of experiments, from 1900 to 2019.[319]

Wikipedia Views

The chart below shows pageviews of the English Wikipedia article Design of experiments, from July 2015 to November 2021.[320]


Meta information on the timeline

How the timeline was built

The initial version of the timeline was written by Sebastian Sanchez.

Funding information for this timeline is available.

Feedback and comments

Feedback for the timeline can be provided at the following places:

  • FIXME

What the timeline is still missing

  • Updated visual data
  • for books: https://academic-accelerator.com/encyclopedia/optimal-design
  • doi: 10.1007/978-3-319-33781-4_1
  • experiment design/design of experiments "in 1800..2020"
  • Add Google Scholar table
  • Vipul: "will this timeline eventually talk of things like double-blinding, triple-blinding, placebos, RCTs, etc., right? You have blinding but I guess the rest are variants on the idea".
  • Vipul: "Cover "Statistical significance", "p-values" and preregistration."
  • Books


Timeline update strategy

See also

References

  1. 1.0 1.1 1.2 "Clinical trials—from ancient Babylon to today". Main Line Health. 17 May 2021. Retrieved 31 March 2026.
  2. Kleisiaris, Christos F.; Sfakianakis, Chrisanthos; Papathanasiou, Ioanna V. (March 15, 2014). "Health care practices in ancient Greece: The Hippocratic ideal". Journal of Medical Ethics and History of Medicine. 7. PMC 4263393. PMID 25512827. Retrieved April 9, 2025.
  3. "Baconian method". Encyclopaedia Britannica. Retrieved April 9, 2025.
  4. Lienhard, John H. (January 4, 1989). "Galileo's Experiment". The Engines of Our Ingenuity. University of Houston. Retrieved April 9, 2025.
  5. "Robert Boyle". Internet Encyclopedia of Philosophy. Retrieved April 9, 2025.
  6. Ore, Øystein (May 1960). "Pascal and the Invention of Probability Theory" (PDF). The American Mathematical Monthly. 67 (5). Mathematical Association of America: 409–419. JSTOR 2309286. Retrieved April 9, 2025.
  7. "Fermat and Pascal on Probability" (PDF). University of York. Retrieved April 9, 2025.
  8. Polasek, Wolfgang (August 2000). "The Bernoullis and the Origin of Probability Theory: Looking back after 300 Years". Resonance – Journal of Science Education. 5 (8): 26–42. Retrieved April 9, 2025.
  9. Cline, Douglas (August 9, 2020). "Age of Enlightenment". Physics LibreTexts. University of Rochester. Retrieved April 9, 2025.
  10. "The 'father of modern statistics' honoured". BBC News. September 9, 2016. Retrieved April 9, 2025.
  11. Schulz, Kathryn (August 14, 2023). "How Carl Linnaeus Set Out to Label All of Life". The New Yorker. Retrieved April 9, 2025.
  12. Schwarz, K. A., & Pfister, R.: Scientific psychology in the 18th century: a historical rediscovery. In: Perspectives on Psychological Science, Nr. 11, p. 399-407.
  13. "Statement on R A Fisher". Rothamsted Research. June 2020. Retrieved April 9, 2025.
  14. 14.00 14.01 14.02 14.03 14.04 14.05 14.06 14.07 14.08 14.09 14.10 14.11 14.12 "1.1 - A Quick History of the Design of Experiments (DOE) | STAT 503". PennState: Statistics Online Courses. Retrieved 11 May 2021.
  15. Preece, D. A. (December 1990). "R. A. Fisher and Experimental Design: A Review". Biometrics. 46 (4). International Biometric Society: 925–935. doi:10.2307/2532438. JSTOR 2532438. Retrieved April 9, 2025.
  16. Biau, David Jean; Jolles, Brigitte M.; Porcher, Raphaël (March 2010). "P Value and the Theory of Hypothesis Testing: An Explanation for New Researchers". Clinical Orthopaedics and Related Research. 468 (3): 885–892. doi:10.1007/s11999-009-1164-4. PMC 2816758. PMID 19921345. Retrieved April 9, 2025.
  17. Kramer, Lloyd; Maza, Sarah (23 June 2006). A Companion to Western Historical Thought. Wiley. ISBN 978-1-4051-4961-7. Shortly after the start of the Cold War [...] double-blind reviews became the norm for conducting scientific medical research, as well as the means by which peers evaluated scholarship, both in science and in history.
  18. 19.0 19.1 19.2 19.3 19.4 19.5 19.6 19.7 "Chapter 2 A Brief History of Experimental Design". JABSTB: Statistical Design and Analysis of Experiments with R. Retrieved 2026-04-07.
  19. 20.00 20.01 20.02 20.03 20.04 20.05 20.06 20.07 20.08 20.09 20.10 20.11 20.12 20.13 20.14 Cite error: Invalid <ref> tag; no text was provided for refs named EncyclopediaExperimentalDesign
  20. 21.0 21.1 21.2 21.3 21.4 21.5 21.6 21.7 21.8 Cite error: Invalid <ref> tag; no text was provided for refs named StudocuBriefHistoryDAE
  21. 22.0 22.1 22.2 22.3 22.4 22.5 22.6 Cite error: Invalid <ref> tag; no text was provided for refs named StudyAndScoreDOE
  22. 24.00 24.01 24.02 24.03 24.04 24.05 24.06 24.07 24.08 24.09 24.10 24.11 24.12 24.13 24.14 24.15 24.16 Cite error: Invalid <ref> tag; no text was provided for refs named ChandramouliDOE100
  23. "Textbooks and other publications on controlled clinical trials". PubMed Central. Retrieved 3 April 2026.
  24. "Design of experiments". Wikipedia. Retrieved 3 April 2026.
  25. "Randomized controlled trial - History". Wikipedia. Retrieved 3 April 2026.
  26. "Randomized controlled trial - applications". Wikipedia. Retrieved 3 April 2026.
  27. "Randomized controlled trial - applications". Wikipedia. Retrieved 3 April 2026.
  28. Gerhard von Rad, Old Testament Theology, Vol. 2 (Louisville: Westminster John Knox, 1965), 307–308.
  29. David C. Lindberg, The Beginnings of Western Science (Chicago: University of Chicago Press, 2007), 14–15.
  30. James C. VanderKam, An Introduction to Early Judaism (Grand Rapids: Eerdmans, 2001), 124.
  31. Hayashi, Takao (2008). "Magic Squares in Indian Mathematics". Encyclopaedia of the History of Science, Technology, and Medicine in Non-Western Cultures (2nd ed.). Springer. pp. 1252–1259. doi:10.1007/978-1-4020-4425-0_9778.
  32. El-Bizri, Nader (2005). "A Philosophical Perspective on Alhazen's Optics". Arabic Sciences and Philosophy. 15 (2): 189–218. doi:10.1017/S0957423905000172.
  33. Aligabi, Zahra (2020). "Reflections on Avicenna's impact on medicine: his reach beyond the Middle East". Journal of Community Hospital Internal Medicine Perspectives. 10 (4): 310–312. doi:10.1080/20009666.2020.1774301. Retrieved 31 March 2026.
  34. Pearce, J. M. S. (29 September 2022). "Francis Bacon's natural philosophy and medicine". Hektoen International: A Journal of Medical Humanities. Retrieved 31 March 2026.
  35. "Ole Rømer Profile: First to Measure the Speed of Light". American Museum of Natural History. Retrieved 13 August 2026.
  36. Boyle, Robert (1683). New Experiments and Observations Touching Cold, or, An Experimental History of Cold, Begun. London: Richard Davis. Retrieved 31 March 2026.
  37. Colbourn, Charles J.; Dinitz, Jeffrey H. Handbook of Combinatorial Designs (2nd ed.). CRC Press. p. 12. ISBN 9781420010541. Retrieved 28 March 2017.
  38. "Choi Seok-jeong (1646–1715)". MacTutor History of Mathematics Archive. University of St Andrews. Retrieved 31 March 2026.
  39. John Arbuthnot (1710). "An argument for Divine Providence, taken from the constant regularity observed in the births of both sexes". Philosophical Transactions of the Royal Society of London. 27 (325–336): 186–190. doi:10.1098/rstl.1710.0011.
  40. 42.0 42.1 Dunn, Peter M. (January 1, 1997). "James Lind (1716-94) of Edinburgh and the treatment of scurvy". Archives of Disease in Childhood: Fetal and Neonatal Edition. 76 (1): F64–5. doi:10.1136/fn.76.1.F64. PMC 1720613. PMID 9059193.
  41. Laplace, P. (1778). "Mémoire sur les probabilités". Mémoires de l'Académie Royale des Sciences de Paris: 227–332.
  42. Euler, Leonhard (1782). "Recherches sur une nouvelle espèce de quarrés magiques" [Investigations into a new type of magic squares]. Verhandelingen Uitgegeven Door Het Zeeuwsch Genootschap der Wetenschappen te Vlissingen (in French). 9: 85–239.{{cite journal}}: CS1 maint: unrecognized language (link)
  43. Donaldson, I. M. L. (December 2005). "Mesmer's 1780 Proposal for a Controlled Trial to Test his Method of Treatment Using 'Animal Magnetism'". Journal of the Royal Society of Medicine. 98 (12): 572–575. doi:10.1177/014107680509801226.
  44. "Kent Academic Repository" (PDF). kar.kent.ac.uk. Retrieved 23 October 2021.
  45. Donaldson, I. M. L. (April 2017). "Antoine de Lavoisier's Role in Designing a Single-Blind Trial to Assess whether 'Animal Magnetism' Exists". Journal of the Royal Society of Medicine. 110 (4): 163–167. doi:10.1177/0141076817699740.
  46. Gould, Stephen J. (1989). "The Chain of Reason vs. The Chain of Thumbs". Natural History. 98 (7): 12–21.
  47. "Carl Friedrich Gauss & Adrien-Marie Legendre Discover the Method of Least Squares". History of Information. Jeremy M. Norman. Retrieved 31 March 2026.
  48. "Polynomial regression". frontend. Retrieved 18 March 2022.
  49. Fétis, François-Joseph (1868). Biographie Universelle des Musiciens et Bibliographie Générale de la Musique, Tome 1 (Second ed.). Paris: Firmin Didot Frères, Fils, et Cie. p. 249. Retrieved 2011-07-21. {{cite book}}: Unknown parameter |name-list-format= ignored (|name-list-style= suggested) (help)
  50. Dubourg, George (1852). The Violin: Some Account of That Leading Instrument and its Most Eminent Professors... (Fourth ed.). London: Robert Cocks and Co. pp. 356–357. Retrieved 2011-07-21. {{cite book}}: Unknown parameter |name-list-format= ignored (|name-list-style= suggested) (help)
  51. Stigler (1986, pp 154–155)
  52. Stolberg, M. (December 2006). "Inventing the randomized double-blind trial: the Nuremberg salt test of 1835". Journal of the Royal Society of Medicine. 99 (12): 642–643. doi:10.1258/jrsm.99.12.642. PMC 1676327. PMID 17139070.
  53. Faerstein, Eduardo; Winkelstein Jr., Warren (September 2012). "Adolphe Quetelet: Statistician and More". Epidemiology. 23 (5): 762–763. doi:10.1097/EDE.0b013e318261c86f. Retrieved 31 March 2026.
  54. Cooper, Max; Middleton, Jo; Cooper, Sarah (2025). "Cold-water, Sulphur and 'the itch': James Henry's principles for conducting controlled trials (1843)". Irish Journal of Medical Science. 194 (6): 2303–2305. doi:10.1007/s11845-025-04027-x. PMC 12769593. PMID 40824558. {{cite journal}}: Check |pmc= value (help); Check |pmid= value (help)
  55. Lindner & Rodger 1997, pg.3
  56. Kirkman, Thomas P. (1847), "On a Problem in Combinations", The Cambridge and Dublin Mathematical Journal, II: 191–204
  57. Steiner, J. (1853), "Combinatorische Aufgabe", Journal für die reine und angewandte Mathematik, 1853 (45): 181–182, doi:10.1515/crll.1853.45.181
  58. Newsom, 2006
  59. 64.0 64.1 "Experimental Psychology: History, Features, and Method". Psychologs Magazine. Retrieved April 9, 2025.
  60. Pasteur, Louis (1861). Sur les corpuscules organisés qui existent dans l'atmosphère: Examen de la doctrine des générations spontanées (in français). Paris: Ch. Lahure et Cie. Retrieved 31 March 2026.
  61. Cavaillon, Jean-Marc; Legout, Sandra (2022). "Louis Pasteur: Between Myth and Reality". Biomolecules. 12 (4): 596. doi:10.3390/biom12040596. PMC 9027159. PMID 35454184. Retrieved 31 March 2026.{{cite journal}}: CS1 maint: unflagged free DOI (link)
  62. Bernard, Claude (2008) [1865]. Introduction à l'étude de la médecine expérimentale. Champs. Paris: Flammarion. ISBN 978-2-08-121793-5.
  63. Peirce, C. S. (August 1967). "Note on the Theory of the Economy of Research". Operations Research. 15 (4): 643–648. doi:10.1287/opre.15.4.643.
  64. Robert Burch (2001). "Charles Sanders Peirce". Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University. Retrieved 6 April 2026.
  65. Peirce, Charles Sanders (1877–1878). "Illustrations of the Logic of Science". Popular Science Monthly. Retrieved 13 August 2026.
  66. Peirce, C. S. (1882), "Introductory Lecture on the Study of Logic" delivered September 1882, published in Johns Hopkins University Circulars, v. 2, n. 19, pp. 11–12, November 1882, see p. 11, Google Books Eprint. Reprinted in Collected Papers v. 7, paragraphs 59–76, see 59, 63, Writings of Charles S. Peirce v. 4, pp. 378–82, see 378, 379, and The Essential Peirce v. 1, pp. 210–14, see 210–1, also lower down on 211.
  67. O'Rourke, M.F. (1992). "Frederick Akbar Mahomed". Hypertension. 19 (2): 212–217. doi:10.1161/01.HYP.19.2.212.
  68. Stigler (1986, pp 314–315)
  69. "Role of the Michelson-Morley experiments in making determinations about competing theories". Archived from the original on 2012-11-07. Retrieved 2003-07-17.
  70. Charles Sanders Peirce and Joseph Jastrow (1885). "On Small Differences in Sensation". Memoirs of the National Academy of Sciences. 3: 73–83. http://psychclassics.yorku.ca/Peirce/small-diffs.htm
  71. Stroebe, W. (2012). The truth about Triplett (1898), but nobody seems to care. Perspectives on Psychological Science, 7, 54-57.
  72. Pearson, Karl (1900). "On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling" (PDF). Philosophical Magazine. Series 5. 50 (302): 157–175. doi:10.1080/14786440009463897.
  73. Nahm, Francis Sahngun (2017). "What the P values really tell us". The Korean Journal of Pain. 30 (4): 241. doi:10.3344/kjp.2017.30.4.241.
  74. Nahm, Francis Sahngun (October 2017). "What the P values really tell us". The Korean Journal of Pain. 30 (4): 241–242. doi:10.3344/kjp.2017.30.4.241. ISSN 2005-9159.
  75. Newman, David H., M.D. (2008). Hippocrates' shadow : secrets from the house of medicine (1st Scribner hardcover ed.). New York, NY: Scribner. ISBN 978-1-4165-5153-9.{{cite book}}: CS1 maint: multiple names: authors list (link)
  76. Samhita, Laasya; Gross, Hans J. (1 November 2013). "The "Clever Hans Phenomenon" revisited". Communicative & Integrative Biology. 6 (6) e27122. doi:10.4161/cib.27122. PMC 3921203. PMID 24563716.
  77. Rivers WH, Webber HN (August 1907). "The action of caffeine on the capacity for muscular work". The Journal of Physiology. 36 (1): 33–47. doi:10.1113/jphysiol.1907.sp001215. PMC 1533733. PMID 16992882.
  78. Gimhan, Vishal (6 March 2025). "Understanding the t-Distribution: A Guide for Small Sample Analysis". Medium. Retrieved 31 March 2026.
  79. Edgeworth, F.Y. (June 1908). "On the Probable Errors of Frequency-Constants". Journal of the Royal Statistical Society. 71 (2): 381–397. JSTOR 2339461.
  80. The Correlation Between Relatives on the Supposition of Mendelian Inheritance. Ronald A. Fisher. Philosophical Transactions of the Royal Society of Edinburgh. 1918. (volume 52, pages 399–433)
  81. Smith, Kirstine (1918). "On the Standard Deviations of Adjusted and Interpolated Values of an Observed Polynomial Function and its Constants and the Guidance they give Towards a Proper Choice of the Distribution of Observations". Biometrika. 12 (1–2). Oxford University Press: 1–85. doi:10.2307/2331929. Retrieved 31 March 2026.
  82. 90.0 90.1 "Experimental Design | Encyclopedia.com". www.encyclopedia.com. Retrieved 5 April 2021.
  83. On the "Probable Error" of a Coefficient of Correlation Deduced from a Small Sample. Ronald A. Fisher. Metron, 1: 3–32 (1921)
  84. Fisher, R.A. (1922). "On the mathematical foundations of theoretical statistics". Philosophical Transactions of the Royal Society of London, Series A. 222 (594–604): 309–368. doi:10.1098/rsta.1922.0009.
  85. Scheffé (1959, p 291, "Randomization models were first formulated by Neyman (1923) for the completely randomized design, by Neyman (1935) for randomized blocks, by Welch (1937) and Pitman (1937) for the Latin square under a certain null hypothesis, and by Kempthorne (1952, 1955) and Wilk (1955) for many other designs.")
  86. Fisher, Ronald A. (1921). "Studies in Crop Variation. I. An Examination of the Yield of Dressed Grain from Broadbalk". Journal of Agricultural Science. 11 (2): 107–135. doi:10.1017/S0021859600003750.
  87. Fisher, Ronald A. (1923). "Studies in Crop Variation. II. The Manurial Response of Different Potato Varieties". Journal of Agricultural Science. 13 (3): 311–320. doi:10.1017/S0021859600003592.
  88. Toews, Ingrid; Anglemyer, Andrew; Nyirenda, John Lz; Alsaid, Dima; Balduzzi, Sara; Grummich, Kathrin; Schwingshackl, Lukas; Bero, Lisa (2024-01-04). "Healthcare outcomes assessed with observational study designs compared with those assessed in randomized trials: a meta-epidemiological study". The Cochrane Database of Systematic Reviews. 1 (1): MR000034. doi:10.1002/14651858.MR000034.pub3. PMID 38174786.
  89. Gosnell, Harold F. (1926). "An Experiment in the Stimulation of Voting". American Political Science Review. 20 (4): 869–874. doi:10.1017/S0003055400110524.
  90. Cumming, Geoff (2011). "From null hypothesis significance to testing effect sizes". Understanding The New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis. Multivariate Applications Series. East Sussex, United Kingdom: Routledge. pp. 21–52. ISBN 978-0-415-87968-2.
  91. Fisher, Ronald A. (1925). Statistical Methods for Research Workers. Edinburgh, UK: Oliver and Boyd. pp. 43. ISBN 978-0-050-02170-5. {{cite book}}: ISBN / Date incompatibility (help)
  92. Lehmann, Erich L. (2011). Fisher, Neyman, and the creation of classical statistics. New York, NY: Springer Science+Business Media, LLC. p. 15. ISBN 978-1-4419-9500-1.
  93. Conniffe, Denis (1990–1991). "R. A. Fisher and the development of statistics—a view in his centenary year". Journal of the Statistical and Social Inquiry Society of Ireland. Vol. XXVI, no. 3. Dublin: Statistical and Social Inquiry Society of Ireland. p. 87. hdl:2262/2764. ISSN 0081-4776.
  94. Savage, Leonard J. (1976). "On Rereading R. A. Fisher". Annals of Statistics. 4 (3): 441–500. doi:10.1214/aos/1176343456.
  95. Kopf, Dan. "An error made in 1925 led to a crisis in modern science—now researchers are joining to fix it". Quartz. Retrieved 13 March 2021.
  96. Cashin, Aidan G.; Hansford, Harrison J.; Hernán, Miguel A.; Swanson, Sonja A.; Lee, Hopin; Jones, Matthew D.; Dahabreh, Issa J.; Dickerman, Barbra A.; Egger, Matthias; Garcia-Albeniz, Xabier; Golub, Robert M.; Islam, Nazrul; Lodi, Sara; Moreno-Betancur, Margarita; Pearson, Sallie-Anne; Schneeweiss, Sebastian; Sharp, Melissa K.; Sterne, Jonathan A. C.; Stuart, Elizabeth A.; McAuley, James H. (2025-09-03). "Transparent Reporting of Observational Studies Emulating a Target Trial-The TARGET Statement". JAMA. 334 (12): 1084–1093. doi:10.1001/jama.2025.13350. PMID 40899949. {{cite journal}}: Check |pmid= value (help)
  97. Box, Joan Fisher (February 1980). "R. A. Fisher and the Design of Experiments, 1922-1926". The American Statistician. 34 (1): 1. doi:10.2307/2682986.
  98. Fisher, Ronald (1926). "The Arrangement of Field Experiments". Journal of the Ministry of Agriculture of Great Britain. 33. London: Ministry of Agriculture and Fisheries: 503–513.
  99. Bernstein, S. N. (1926). "Sur l'extension du théorème limite du calcul des probabilités aux sommes de quantités dépendantes". Mathematische Annalen. 97: 1–59.
  100. Neyman, J; Pearson, E. S. (1 January 1933). "On the Problem of the most Efficient Tests of Statistical Hypotheses". Philosophical Transactions of the Royal Society A. 231 (694–706): 289–337. Bibcode:1933RSPTA.231..289N. doi:10.1098/rsta.1933.0009.
  101. Fisher, R.A. (1935). The Design of Experiments. Oliver and Boyd. pp. 114–145.
  102. Box, JF (February 1980). "R. A. Fisher and the Design of Experiments, 1922–1926". The American Statistician. 34 (1): 1–7. doi:10.2307/2682986. JSTOR 2682986.
  103. Yates, F (June 1964). "Sir Ronald Fisher and the Design of Experiments". Biometrics. 20 (2): 307–321. doi:10.2307/2528399. JSTOR 2528399.
  104. Stanley, Julian C. (1966). "The Influence of Fisher's "The Design of Experiments" on Educational Research Thirty Years Later". American Educational Research Journal. 3 (3): 223–229. doi:10.3102/00028312003003223. JSTOR 1161806.
  105. Fisher, Ronald A. (1971) [1935]. The Design of Experiments (9th ed.). Macmillan. ISBN 0-02-844690-9.
  106. Box, Joan Fisher (1978). R.A. Fisher, The Life of a Scientist. New York: Wiley. p. 134. ISBN 0-471-09300-9.
  107. "Earliest Known Uses of Some of the Words of Mathematics (F)". jeff560.tripod.com. Retrieved 13 August 2026.
  108. Crawford, Meredith P. (1937). The Coöperative Solving of Problems by Young Chimpanzees. Johns Hopkins Press.
  109. Bose, R. C.; Nair, K. R. (1939), "Partially balanced incomplete block designs", Sankhyā, 4: 337–372
  110. Fisher, R.A. (1940), "An examination of the different possible solutions of a problem in incomplete blocks", Annals of Eugenics, 10: 52–75, doi:10.1111/j.1469-1809.1940.tb02237.x, hdl:2440/15239
  111. Kishen, K. (1942), "On latin and hyper-graeco cubes and hypercubes", Current Science, 11: 98–99
  112. Chalmers I, Clarke M (April 2004). "Commentary: the 1944 patulin trial: the first properly controlled multicentre trial conducted under the aegis of the British Medical Research Council". International Journal of Epidemiology. 33 (2): 253–260. doi:10.1093/ije/dyh162. PMID 15082623.
  113. Wald, Abraham (1945). "Sequential Tests of Statistical Hypotheses". Annals of Mathematical Statistics. 16 (2): 117–186. doi:10.1214/aoms/1177731118.
  114. Finney, D. J. (1945). "The fractional replication of factorial arrangements". Annals of Eugenics. 12: 291–301. doi:10.1111/j.1469-1809.1943.tb02333.x.
  115. National Research Council (1995). Statistical Methods for Testing and Evaluating Defense Systems: Interim Report (Report). Washington, D.C.: The National Academies Press.
  116. Jellinek, E. M. "Clinical Tests on Comparative Effectiveness of Analgesic Drugs", Biometrics Bulletin, Vol.2, No.5, (October 1946), pp.87–91.
  117. "5.3.3.5. Plackett-Burman designs". www.itl.nist.gov. Retrieved 22 July 2023.
  118. Rao, C.R. (1946), "Hypercubes of strength d leading to confounded designs in factorial experiments", Bulletin of the Calcutta Mathematical Society, 38: 67–78
  119. Rao, C.R. (1947), "Factorial experiments derivable from combinatorial arrangements of arrays", Supplement to the Journal of the Royal Statistical Society, 9 (1): 128–139, doi:10.2307/2983576, JSTOR 2983576
  120. Metcalfe, N.H. (2011). "Sir Geoffrey Marshall (1887-1982): respiratory physician, catalyst for anaesthesia development, doctor to both Prime Minister and King, and World War I Barge Commander". Journal of Medical Biography. 19 (1): 10–14. doi:10.1258/jmb.2010.010019. PMID 21350072.
  121. "The Nuremberg Code". U.S. Department of Health & Human Services. Retrieved 13 August 2026.
  122. "Nuremberg Code". The Doctor's Trial: The Medical Case of the Subsequent Nuremberg Proceedings. United States Holocaust Memorial Museum Online Exhibitions. Retrieved 13 February 2019.
  123. Shuster, Evelyne (1997). "Fifty Years Later: The Significance of the Nuremberg Code". New England Journal of Medicine. 337 (20): 1436–1440. doi:10.1056/NEJM199711133372006. PMID 9358142.
  124. Anscombe, F. J. (1948). "The Validity of Comparative Experiments". Journal of the Royal Statistical Society. Series A (General). 111 (3): 181–211. doi:10.2307/2984159. JSTOR 2984159.
  125. "The MRC randomized trial of streptomycin and its legacy". PubMed Central. Retrieved 3 April 2026.
  126. Healy, M. J. R. (1995). "Frank Yates, 1902-1994: The Work of a Statistician". International Statistical Review / Revue Internationale de Statistique. 63 (3): 271–288. ISSN 0306-7734.
  127. Grundy, P. M.; Healy, M. J. R. (1950). "Restricted Randomization and Quasi-Latin Squares". Journal of the Royal Statistical Society. Series B (Methodological). 12 (2): 286–291. ISSN 0035-9246.
  128. Navarro, Mario; Siegel, Jason T. (2018). "Solomon Four-Group Design". SAGE Publications. Retrieved 22 November 2019.
  129. Template:Cite thesis
  130. Kenneth J. Arrow, David Blackwell and M.A. Girshick (1949). "Bayes and minimax solutions of sequential decision problems". Econometrica. 17 (3/4): 213–244. JSTOR 1905525.
  131. Bruck, R.H.; Ryser, H.J. (1949), "The nonexistence of certain finite projective planes", Canadian Journal of Mathematics, 1: 88–93, doi:10.4153/cjm-1949-009-2
  132. Grundy, P.M.; Healy, M.J.R. (1950). "Restricted randomization and quasi-Latin squares". Journal of the Royal Statistical Society, Series B. 12 (2): 286–291. doi:10.1111/j.2517-6161.1950.tb00062.x.
  133. Template:Cite thesis
  134. Doll R, Hill AB (1950). "Smoking and carcinoma of the lung; preliminary report". British Medical Journal. 2 (4682): 739–748. doi:10.1136/bmj.2.4682.739. PMC 2038856. PMID 14772469.
  135. Chowla, S.; Ryser, H.J. (1950), "Combinatorial problems", Canadian Journal of Mathematics, 2: 93–99, doi:10.4153/cjm-1950-009-8
  136. Draper, Norman R. (1992). "Introduction to Box and Wilson (1951) On the Experimental Attainment of Optimum Conditions". Breakthroughs in Statistics: Methodology and Distribution. Springer. pp. 267–269. doi:10.1007/978-1-4612-4380-9_22.
  137. Martins, Joaquim R. R. A.; Ning, Andrew (January 2022). "A Short History of Optimization". Engineering Design Optimization. Cambridge University Press. doi:10.1017/9781108980647.
  138. Robbins, Herbert (1952). "Some aspects of the sequential design of experiments". Bulletin of the American Mathematical Society. 58 (5): 527–535. doi:10.1090/S0002-9904-1952-09620-8.
  139. Horvitz, D. G.; Thompson, D. J. (1952). "A generalization of sampling without replacement from a finite universe". Journal of the American Statistical Association. 47 (260): 663–685. doi:10.1080/01621459.1952.10483446. JSTOR 2280784.
  140. Bose, R. C.; Shimamoto, T. (June 1952). "Classification and Analysis of Partially Balanced Incomplete Block Designs with Two Associate Classes". Journal of the American Statistical Association. 47 (258): 151–184. doi:10.1080/01621459.1952.10501161.
  141. Hróbjartsson A, Gøtzsche PC (May 2001). "Is the placebo powerless? An analysis of clinical trials comparing placebo with no treatment". The New England Journal of Medicine. 344 (21): 1594–602. doi:10.1056/NEJM200105243442106. PMID 11372012.
  142. Wilk, M.B. (1955). "The Randomization Analysis of a Generalized Randomized Block Design". Biometrika. 42 (1–2): 70–79. doi:10.2307/2333423.
  143. Lindley, D. V. (1956). "On a measure of information provided by an experiment". Annals of Mathematical Statistics. 27 (4): 986–1005. doi:10.1214/aoms/1177728069.
  144. Doll R, Hill AB (1956). "Lung cancer and other causes of death in relation to smoking; a second report on the mortality of British doctors". British Medical Journal. 2 (5001): 1071–1081. doi:10.1136/bmj.2.5001.1071. PMC 2035864. PMID 13364389.
  145. Kish, L. (1959). "Some statistical problems in research design". American Sociological Review. 26 (3): 328–338. doi:10.2307/2089381.
  146. Pierson, Raymond H.; Fay, Edward A. (December 1959). "Guidelines for Interlaboratory Testing Programs". Analytical Chemistry. 31 (12): 25A – 49A. doi:10.1021/ac60156a708.
  147. Bose, R. C.; Mesner, D. M. (1959). "On linear associative algebras corresponding to association schemes of partially balanced designs". Annals of Mathematical Statistics. 30 (1): 21–38. doi:10.1214/aoms/1177706356.
  148. Russell, W.M.S.; Burch, R.L. (1959). The Principles of Humane Experimental Technique. London: Methuen. ISBN 0-900767-78-2. {{cite book}}: ISBN / Date incompatibility (help)
  149. "Genichi Taguchi". asq.org. American Society for Quality. Retrieved 31 March 2026.
  150. Box, George E. P.; Behnken, Donald (1960). "Some new three level designs for the study of quantitative variables". Technometrics. 2: 455–475. doi:10.1080/00401706.1960.10489912.
  151. Ranade, Shruti Sunil; Thiagarajan, Padma (November 2017). "Selection of a design for response surface". IOP Conference Series: Materials Science and Engineering. 263: 022043. doi:10.1088/1757-899X/263/2/022043.
  152. Wason, Peter C. (1960). "On the failure to eliminate hypotheses in a conceptual task". Quarterly Journal of Experimental Psychology. 12 (3): 129–140. doi:10.1080/17470216008416717.
  153. Thistlethwaite, D.; Campbell, D. (1960). "Regression-Discontinuity Analysis: An alternative to the ex post facto experiment". Journal of Educational Psychology. 51 (6): 309–317. doi:10.1037/h0044319.
  154. Box, G. E. P.; Hunter, J. S. (1961). "The 2^(k-p) fractional factorial designs". Technometrics. 3: 311–351.
  155. "Definition of NOCEBO". www.merriam-webster.com. Retrieved 5 March 2022.
  156. Kennedy, 1961
  157. Harman WW, McKim RH, Mogar RE, Fadiman J, Stolaroff MJ (August 1966). "Psychedelic agents in creative problem-solving: a pilot study". Psychological Reports. 19 (1): 211–227. doi:10.2466/pr0.1966.19.1.211. PMID 5942087.
  158. Kôno, Kazumasa (1962). "Optimum designs for quadratic regression on k-cube". Memoirs of the Faculty of Science. Kyushu University. Series A. Mathematics. 16 (2): 114–122. doi:10.2206/kyushumfs.16.114.
  159. Smith, Vernon L. (1962). "An Experimental Study of Competitive Market Behavior". Chapman University Digital Commons. Chapman University. Retrieved 31 March 2026.
  160. Stankova, Tatiana (30 June 2020). "Application of Nelder wheel experimental design in forestry research". Silva Balcanica. 21 (1): 29–40. doi:10.3897/silvabalcanica.21.e54425.{{cite journal}}: CS1 maint: unflagged free DOI (link)
  161. 176.0 176.1 Heppner, Puncky Paul; Wampold, Bruce E.; Owen, Jesse; Wang, Kenneth T. (21 August 2015). Research Design in Counseling. Cengage Learning. ISBN 978-1-305-46501-5.
  162. Boren, J. J. (1963). "Repeated acquisition of new behavioral chains". American Psychologist, 18, p. 421.
  163. Kish, Leslie (1965). Survey Sampling. New York: John Wiley & Sons, Inc. ISBN 0-471-10949-5.
  164. Selvin, H.C.; Stuart, A. (1966). "Data-Dredging Procedures in Survey Analysis". The American Statistician. 20 (3): 20–23. doi:10.1080/00031305.1966.10480401.
  165. Rosenthal, R. (1966). Experimenter Effects in Behavioral Research. New York: Appleton-Century-Crofts.
  166. Myers, Raymond H. (1971). Response Surface Methodology. Boston: Allyn and Bacon.
  167. Chernoff, H. (1972) Sequential Analysis and Optimal Design, SIAM Monograph.
  168. Delsarte, P. (1973), "An Algebraic Approach to the Association Schemes of Coding Theory", Philips Research Reports (Supplement No. 10)
  169. Rubin, Donald B. (1973). "Matching to Remove Bias in Observational Studies". Biometrics. 29 (1): 159–183. doi:10.2307/2529684. JSTOR 2529684.
  170. Rubin, Donald B. (1974). "Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies". Journal of Educational Psychology. 66 (5): 688–701. doi:10.1037/h0037350.
  171. Denniston, R. H. F. (September 1974). "Denniston's paper, open access". Discrete Mathematics. 9 (3): 229–233. doi:10.1016/0012-365X(74)90004-1.
  172. Ray-Chaudhuri, Dijen K.; Wilson, Richard M. (1975), "On t-designs", Osaka Journal of Mathematics, 12: 737–744
  173. Atkinson, A. C.; Fedorov, V. V. (1975). "The design of experiments for discriminating between two rival models". Biometrika. 62 (1): 57–70. doi:10.1093/biomet/62.1.57.
  174. Pocock, Stuart J.; Simon, Richard (Mar 1975). "Sequential Treatment Assignment with Balancing for Prognostic Factors in the Controlled Clinical Trial". Biometrics. 31 (1). International Biometric Society: 103–115. doi:10.2307/2529712. JSTOR 2529712. PMID 1100130.
  175. James Scheirer, William S. Ray, Nathan Hare: "The Analysis of Ranked Data Derived from Completely Randomized Factorial Designs." Biometrics 32(2), 1976, pp. 429–434, Template:Doi
  176. Montgomery, Douglas C. (2013). Design and Analysis of Experiments. John Wiley & Sons Incorporated. ISBN 978-1-62198-227-2.
  177. Miettinen, O. (1976). "Estimability and estimation in case–referent studies". American Journal of Epidemiology. 103 (2): 226–235. doi:10.1093/oxfordjournals.aje.a112220. PMID 1251836.
  178. James Scheirer, William S. Ray, Nathan Hare: "The Analysis of Ranked Data Derived from Completely Randomized Factorial Designs." Biometrics 32(2), 1976, pp. 429–434, Template:Doi
  179. Pocock S (2005). "When (not) to stop a clinical trial for benefit" (PDF). JAMA. 294 (17): 2228–2230. doi:10.1001/jama.294.17.2228. PMID 16264167.
  180. Eglajs, V.; Audze P. (1977). "New approach to the design of multifactor experiments". Problems of Dynamics and Strengths. 35 (in Russian). Riga: Zinatne Publishing House: 104–107.{{cite journal}}: CS1 maint: unrecognized language (link)
  181. Pocock, Stuart J. (August 1977). "Group Sequential Methods in the Design and Analysis of Clinical Trials". Biometrika. 64 (2): 191–199. doi:10.2307/2335684.
  182. Rubin, Donald (1978). "Bayesian Inference for Causal Effects: The Role of Randomization". The Annals of Statistics. 6 (1): 34–58. doi:10.1214/aos/1176344064.
  183. "Experimental Design - an overview | ScienceDirect Topics". www.sciencedirect.com. Retrieved 23 March 2021.
  184. Ma, Will. "Lecture 1: Four (and a Half) Proofs of the Basic Prophet Inequality" (PDF). Columbia University. Columbia Business School. Retrieved 31 March 2026.
  185. Richter, Felicitas; Dewey, Marc (September 2014). "Zelen Design in Randomized Controlled Clinical Trials". Radiology. 272 (3): 919–919. doi:10.1148/radiol.14140834.
  186. Homer, Caroline S.E. (April 2002). "Using the Zelen design in randomized controlled trials: debates and controversies". Journal of Advanced Nursing. 38 (2): 200–207. doi:10.1046/j.1365-2648.2002.02164.x.
  187. Cook, Thomas D. and Donald T. Campbell (1979), Quasi-experimentation: Design & Analysis Issues for Field Settings. Boston: Houghton-Mifflin
  188. O'Brien, Peter C.; Fleming, Thomas R. (1979). "A Multiple Testing Procedure for Clinical Trials". Biometrics. 35 (3): 549–556. doi:10.2307/2530245. JSTOR 2530245. PMID 497341.
  189. McKay, M.D.; Beckman, R.J.; Conover, W.J. (May 1979). "A Comparison of Three Methods for Selecting Values of Input Variables in the Analysis of Output from a Computer Code". Technometrics. 21 (2). American Statistical Association: 239–245. doi:10.2307/1268522. ISSN 0040-1706. JSTOR 1268522.
  190. Gittins, J.C. (1979). "Bandit Processes and Dynamic Allocation Indices". Journal of the Royal Statistical Society, Series B. 41 (2): 148–177. doi:10.1111/j.2517-6161.1979.tb01068.x.
  191. Coronary Drug Project Research Group (October 1980). "Influence of adherence to treatment and response of cholesterol on mortality in the coronary drug project". New England Journal of Medicine. 303 (18): 1038–1041. doi:10.1056/NEJM198010303031804. PMID 6999345.
  192. Tukey, John W. (1980). "We Need Both Exploratory and Confirmatory". The American Statistician. 34 (1): 23–25. doi:10.2307/2682991. JSTOR 2682991.
  193. Karkar, Ravi; Zia, Jasmine; Vilardaga, Roger; Mishra, Sonali R; Fogarty, James; Munson, Sean A; Kientz, Julie A (1 May 2016). "A framework for self-experimentation in personalized health". Journal of the American Medical Informatics Association. 23 (3): 440–448. doi:10.1093/jamia/ocv150.
  194. George E.P., Box (2006). Improving Almost Anything: Ideas and Essays (Revised ed.). Hoboken, New Jersey: Wiley.
  195. Rubin, Donald B.; Rosenbaum, Paul R. (1983). "The Central Role of the Propensity Score in Observational Studies for Causal Effects". Biometrika. 70 (1): 41–55. doi:10.2307/2335942.
  196. Lan, K. K. Gordon; DeMets, David L. (1983). "Discrete Sequential Boundaries for Clinical Trials". Biometrika. 70 (3): 659–663. doi:10.2307/2336502. JSTOR 2336502.
  197. Begaud B (1984). "Standardized assessment of adverse drug reactions: the method used in France. Special workshop—clinical". Drug Information Journal. 18 (3–4): 275–281. doi:10.1177/009286158401800314.
  198. "Revisiting Hurlbert 1984". Reflections on Papers Past. 29 November 2020. Retrieved 29 March 2022.
  199. LaLonde, Robert (1986). "Evaluating the Econometric Evaluations of Training Programs with Experimental Data". American Economic Review. 4 (76): 604–620.
  200. Street, Anne Penfold; Street, Professor of Mathematics Anne Penfold; Street, Deborah J.; Street, Lecturer in Biometry Deborah J. (1987). "Combinatorics of Experimental Design". Clarendon Press.
  201. Wang, S. K.; Tsiatis, A. A. (1987). "Approximately optimal one-parameter boundaries for group sequential trials". Biometrics. 43 (1): 193–199. doi:10.2307/2531959. ISSN 0006-341X. JSTOR 2531959. PMID 3567304.
  202. Street, Anne Penfold; Street, Professor of Mathematics Anne Penfold; Street, Deborah J.; Street, Lecturer in Biometry Deborah J. (1987). "Combinatorics of Experimental Design". Clarendon Press.
  203. Street, Anne Penfold; Street, Deborah J. (1987). "Combinatorics of Experimental Design". Clarendon Press.
  204. Hall, Marshall Jr. (January–February 1989), "Review of Combinatorics of Experimental Design", American Scientist, 77 (1): 91, JSTOR 27855619
  205. Mead, R. (26 July 1990). The Design of Experiments: Statistical Principles for Practical Applications. Cambridge University Press. ISBN 978-0-521-28762-3.
  206. Latham, Gary P.; Erez, Miriam; Locke, Edwin A. (1988). "Resolving scientific disputes by the joint design of crucial experiments by the antagonists: Application to the Erez–Latham dispute regarding participation in goal setting". Journal of Applied Psychology. 73 (4): 753–772. doi:10.1037/0021-9010.73.4.753.
  207. He, Li (17 July 2003). "Design of Experiments Software, DOE software". The Chemical Information Network.
  208. Haaland, Perry D. (1989). Experimental design in biotechnology. New York: Marcel Dekker. ISBN 9780824778811.
  209. Haaland, Perry D. (June 1991). "BOOK REVIEW: EXPERIMENTAL DESIGN IN BIOTECHNOLOGY Perry D. Haaland Marcel Dekkwe, Inc., New York, 1989". Drying Technology. 9 (3): 817–817. doi:10.1080/07373939108916715.
  210. Haaland, Perry D. (25 November 2020). "Experimental Design in Biotechnology". doi:10.1201/9781003065968. {{cite journal}}: Cite journal requires |journal= (help)
  211. Sacks, Jerome; Welch, William J.; Mitchell, Toby J.; Wynn, Henry P. (1989). "Design and Analysis of Computer Experiments". Statistical Science. 4 (4): 409–423. doi:10.1214/ss/1177012413.
  212. "Read "Statistical Methods for Testing and Evaluating Defense Systems: Interim Report" at NAP.edu". Retrieved 14 March 2021.
  213. Lam, C. W. H. (1991), "The Search for a Finite Projective Plane of Order 10", American Mathematical Monthly, 98 (4): 305–318, doi:10.2307/2323798, JSTOR 2323798
  214. Browne, Malcolm W. (20 December 1988), "Is a Math Proof a Proof If No One Can Check It?", The New York Times
  215. Horne, G., & Schwierz, K. (2008). Data farming around the world overview. Paper presented at the 1442-1447. doi:10.1109/WSC.2008.4736222
  216. Pearl, J. (1993). "Aspects of Graphical Models Connected With Causality". Proceedings of the 49th Session of the International Statistical Science Institute. pp. 391–401.
  217. Breggin, Ginger Ross; Breggin, Peter Roger (1995). Talking back to Prozac: what doctors won't tell you about today's most controversial drug. New York: St. Martin's Paperbacks. ISBN 978-0-312-95606-6.
  218. Card, David; Krueger, Alan B. (1994). "Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania". American Economic Review. 84 (4): 772–793. JSTOR 2118030.
  219. Neyer, Barry T. (February 1994). "A D-Optimality-Based Sensitivity Test". Technometrics. 36 (1): 61. doi:10.2307/1269199.
  220. Kish, Leslie (1995). "Methods for design effects". Journal of Official Statistics. 11 (1): 55.
  221. Chaloner, Kathryn; Verdinelli, Isabella (1995), "Bayesian experimental design: a review", Statistical Science, 10 (3): 273–304, doi:10.1214/ss/1177009939
  222. "Integrated Addendum to ICH E6(R1): Guideline for Good Clinical Practice E6(R2)" (PDF). International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. 9 November 2016. Retrieved 31 March 2026.
  223. Jadad, A.R.; Moore R.A.; Carroll D.; Jenkinson C.; Reynolds D.J.M.; Gavaghan D.J.; McQuay H.J. (1996). "Assessing the quality of reports of randomized clinical trials: Is blinding necessary?". Controlled Clinical Trials. 17 (1): 1–12. doi:10.1016/0197-2456(95)00134-4. PMID 8721797.
  224. He, Li (17 July 2003). "Design of Experiments Software, DOE software". The Chemical Information Network.
  225. Begg, C.; Cho, M.; Eastwood, S.; Horton, R.; Moher, D.; Olkin, I.; Pitkin, R.; Rennie, D.; Schulz, K.F.; Simel, D.; Stroup, D.F. (1996). "Improving the quality of reporting of randomized controlled trials: the CONSORT statement". JAMA. 276 (8): 637–639.
  226. Mitchell, T. (1997). Machine Learning. New York: McGraw Hill. ISBN 0-07-042807-7.
  227. Shalala, D.E. (26 August 1998). "Policy statement on changing the population standard used for age adjusting death rates in DHHS publications". Retrieved 13 August 2026.
  228. Brandstein, A.; Horne, G. (1998). "Data Farming: A Meta-Technique for Research in the 21st Century". Maneuver Warfare Science. Quantico, VA: Marine Corps Combat Development Command.
  229. Dehejia, R.H.; Wahba, S. (1999). "Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs". Journal of the American Statistical Association. 94 (448): 1053–1062.
  230. "Analyzing Families of Experiments in SE: a Systematic Mapping Study" (PDF). arxiv.org. Retrieved 12 March 2022.
  231. Oehlert, Gary W. (2010). A First Course in Design and Analysis of Experiments. Gary W. Oehlert.
  232. Ahmad, O.B.; Boschi-Pinto, C.; Lopez, A.D.; Murray, C.J.L.; Lozano, R.; Inoue, M. (2001). "Age Standardization of Rates: A New WHO Standard" (PDF). GPE Discussion Paper Series: No.31. World Health Organization.
  233. "Adversarial Collaboration: An EDGE Lecture by Daniel Kahneman | Edge.org". www.edge.org. Retrieved 8 March 2022.
  234. Berger, Michele W.; (University of Pennsylvania). "In the pursuit of scientific truth, working with adversaries can pay off". phys.org. Retrieved 13 August 2026.
  235. Moher, D.; Schulz, K.F.; Altman, D.G. (2001). "The CONSORT statement: revised recommendations for improving the quality of reports of parallel-group randomized trials". Annals of Internal Medicine. 134 (8): 657–662.
  236. "WMA - Policy". Archived from the original on 20 February 2009.
  237. Musch, J., & Klauer, K. C. (2002). Psychological experimenting on the World Wide Web: Investigating content effects in syllogistic reasoning. In B. Batinic, U.-D. Reips, & M. Bosnjak (Eds.), Online social sciences (pp. 181–212). Hogrefe & Huber Publishers.
  238. Reips, U.-D. (2002). Standards for internet-based experimenting. Experimental Psychology, 49(4), 243–256.
  239. Bloom, H.S.; Michalopoulos, C.; Hill, C.J.; Lei, Y. (2002). "Can Nonexperimental Comparison Group Methods Match the Findings from a Random Assignment Evaluation of Mandatory Welfare-to-Work Programs?". MDRC Working Papers on Research Methodology. {{cite web}}: Missing or empty |url= (help)
  240. Schneider, ed. by Sandra L.; Shanteau, James (2003). Emerging perspectives on judgment and decision research. Cambridge [u.a.]: Cambridge Univ. Press. pp. 438–9. ISBN 052152718X. {{cite book}}: |first= has generic name (help)
  241. "Randomized controlled trials in development economics". Wikipedia. Retrieved 3 April 2026.
  242. Hirata, S. (2003). "Cooperation in chimpanzees". Hattatsu. 95: 103–111.
  243. Campbell MK, Elbourne DR, Altman DG (2004). "CONSORT statement: extension to cluster randomised trials". BMJ. 328 (7441): 702–708. doi:10.1136/bmj.328.7441.702. PMC 381234. PMID 15031246.
  244. Pildal J, Chan AW, Hróbjartsson A, Forfang E, Altman DG, Gøtzsche PC (2005). "Comparison of descriptions of allocation concealment in trial protocols and the published reports: cohort study". BMJ. 330 (7499): 1049. doi:10.1136/bmj.38414.422650.8F. PMC 557221. PMID 15817527.
  245. Pocock S (2005). "When (not) to stop a clinical trial for benefit" (PDF). JAMA. 294 (17): 2228–2230. doi:10.1001/jama.294.17.2228. PMID 16264167.
  246. Ioannidis, John P. A. (2005-08-30). "Why Most Published Research Findings Are False". PLOS Medicine. 2 (8): e124. doi:10.1371/journal.pmed.0020124. PMC 1182327. PMID 16060722.
  247. Kleijnen, J.P.C.; Sanchez, S.M.; Lucas, T.W.; Cioppa, T.M. (2005). "A User's Guide to the Brave New World of Designing Simulation Experiments". INFORMS Journal on Computing. 17 (3): 263–289. doi:10.1287/ijoc.1050.0136.
  248. Swan M (June 2013). "The Quantified Self: Fundamental Disruption in Big Data Science and Biological Discovery". Big Data. 1 (2): 85–99. doi:10.1089/big.2012.0002. PMID 27442063.
  249. "Exploratory IND Studies, Guidance for Industry, Investigators, and Reviewers" (PDF). Food and Drug Administration. January 2006.
  250. Melis, Alicia P.; Hare, Brian; Tomasello, Michael (2006). "Engineering cooperation in chimpanzees: Tolerance constraints on cooperation". Animal Behaviour. 72 (2): 275–286. doi:10.1016/j.anbehav.2005.09.018.
  251. Wisely, Janet (2007). "Building on Improvement: Establishing a National Research Ethics Service". Research Ethics. 3: 3–4. doi:10.1177/174701610700300102. S2CID 167972296.
  252. McCrary (2008). "Manipulation of the Running Variable in the Regression Discontinuity Design: A Density Test". Journal of Econometrics. 142 (2): 698–714. doi:10.1016/j.jeconom.2007.05.005.
  253. Wood, L; Egger, M; Gluud, LL; Schulz, KF; Jüni, P; Altman, DG; Gluud, C; Martin, RM; Wood, AJ; Sterne, JA (2008). "Empirical evidence of bias in treatment effect estimates in controlled trials with different interventions and outcomes: meta-epidemiological study". BMJ. 336 (7644): 601–605. doi:10.1136/bmj.39465.451748.AD. PMC 2267990. PMID 18316340.
  254. Kahneman, Daniel; Klein, Gary. Conditions for intuitive expertise: A failure to disagree. American Psychologist, Vol 64(6), Sep 2009, 515-526. doi: 10.1037/a0016755
  255. Wagenmakers, E.-J., Wetzels, R., Borsboom, D., & van der Maas, H. L. J. (2010). Why psychologists must change the way they analyze their data: The case of psi.
  256. Hróbjartsson A, Gøtzsche PC (January 2010). Hróbjartsson A (ed.). "Placebo interventions for all clinical conditions" (PDF). The Cochrane Database of Systematic Reviews. 106 (1): CD003974. doi:10.1002/14651858.CD003974.pub3. PMID 20091554.
  257. Kilkenny, Carol; Browne, William J.; Cuthill, Innes C.; Emerson, Michael; Altman, Douglas G. (29 June 2010). "Improving Bioscience Research Reporting: The ARRIVE Guidelines for Reporting Animal Research". PLOS Biology. 8 (6) e1000412. doi:10.1371/journal.pbio.1000412. PMC 2893951. PMID 20613859.{{cite journal}}: CS1 maint: unflagged free DOI (link)
  258. Kilkenny, Carol; Parsons, Nick; Kadyszewski, Ed; Festing, Michael F. W.; Cuthill, Innes C.; Fry, Derek; Hutton, Jane; Altman, Douglas G. (30 November 2009). "Survey of the Quality of Experimental Design, Statistical Analysis and Reporting of Research Using Animals". PLOS ONE. 4 (11) e7824. doi:10.1371/journal.pone.0007824. PMC 2779358. PMID 19956596.{{cite journal}}: CS1 maint: unflagged free DOI (link)
  259. Center for Drug Evaluation and Research (CDER); Center for Biologics Evaluation and Research (CBER) (February 2010). "Adaptive Design Clinical Trials for Drugs and Biologics" (PDF). U.S. Food and Drug Administration.
  260. Malone, J.; Holloway, E.; Adamusiak, T.; Kapushesky, M.; Zheng, J.; Kolesnikov, N.; Zhukova, A.; Brazma, A.; Parkinson, H. (2010). "Modeling sample variables with an Experimental Factor Ontology". Bioinformatics. 26 (8): 1112–1118. doi:10.1093/bioinformatics/btq099. PMC 2853691. PMID 20200009.
  261. Schulz, K.F.; Altman, D.G.; Moher, D. (2010). "CONSORT 2010 statement: updated guidelines for reporting parallel group randomised trials". BMJ. 340 c332.
  262. Iacus, Stefano M.; King, Gary; Porro, Giuseppe (2011). "Multivariate Matching Methods That Are Monotonic Imbalance Bounding". Journal of the American Statistical Association. 106 (493): 345–361. doi:10.1198/jasa.2011.tm09599. hdl:2434/151476.
  263. Wang, Shirley S. (30 December 2013). "Health: Scientists Look to Improve Cost and Time of Drug Trials". Wall Street Journal. Retrieved 13 August 2026.
  264. Deng, Alex; Xu, Ya; Kohavi, Ron; Walker, Toby (2013). "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data". WSDM 2013: Sixth ACM International Conference on Web Search and Data Mining.
  265. Simonsohn, Uri; Nelson, Leif D.; Simmons, Joseph P. (2014). "P-curve: A key to the file-drawer". Journal of Experimental Psychology: General. 143 (2): 534–547. doi:10.1037/a0033242. PMID 23855496.
  266. Template:Cite arXiv
  267. Bohannon, John (27 May 2015). "I Fooled Millions Into Thinking Chocolate Helps Weight Loss. Here's How". Gizmodo. Retrieved 13 August 2026.
  268. Head, Megan L.; Holman, Luke; Lanfear, Rob; Kahn, Andrew T.; Jennions, Michael D. (2015-03-13). "The Extent and Consequences of P-Hacking in Science". PLOS Biology. 13 (3) e1002106. doi:10.1371/journal.pbio.1002106. PMC 4359000. PMID 25768323.
  269. Miller, Claire Cain (25 February 2016). "Is Blind Hiring the Best Hiring?". The New York Times. Retrieved 13 August 2026.
  270. Athey, Susan, and Guido Imbens (2016), "Recursive partitioning for heterogeneous causal effects." Proceedings of the National Academy of Sciences 113, (27), 7353–7360.
  271. "International Collaborative Network for N-of-1 Trials and Single-Case Designs". N-of-1 and SCED. Retrieved 2024-07-09.
  272. Aronow, Peter M.; Samii, Cyrus (2017-12-01). "Estimating average causal effects under general interference, with application to a social network experiment". The Annals of Applied Statistics. 11 (4): 1912–1947. arXiv:1305.6156. doi:10.1214/16-aoas1005.
  273. Kennedy, Andrew D. M.; Torgerson, David J.; Campbell, Marion K.; Grant, Adrian M. (2017). "Subversion of allocation concealment in a randomised controlled trial: a historical case study". Trials. 18 (1): 204. doi:10.1186/s13063-017-1946-z. PMC 5414185. PMID 28464922.{{cite journal}}: CS1 maint: unflagged free DOI (link)
  274. Gordon, Brett R.; Zettelmeyer, Florian; Bhargava, Neha; Chapsky, Dan (2018). "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook".
  275. Nosek, Brian A.; Ebersole, Charles R.; DeHaven, Alexander C.; Mellor, David T. (2018-03-13). "The preregistration revolution". Proceedings of the National Academy of Sciences. 115 (11): 2600–2606. doi:10.1073/pnas.1708274114.
  276. Schulz, KF; Chalmers, I; Altman, DG; Grimes, DA; Moher, D; Hayes, RJ (June 2018). "'Allocation concealment': the evolution and adoption of a methodological term". Journal of the Royal Society of Medicine. 111 (6): 216–224. doi:10.1177/0141076818776604. PMC 6022887. PMID 29877772.
  277. "Nobel Prize 2019 experimental approach to poverty". Wikipedia. Retrieved 3 April 2026.
  278. "Adaptive designs for clinical trials of drugs and biologics: Guidance for industry". U.S. Food and Drug Administration (FDA). 1 November 2019. Retrieved 7 April 2021.
  279. King, Gary; Nielsen, Richard (October 2019). "Why Propensity Scores Should Not Be Used for Matching". Political Analysis. 27 (4): 435–454. doi:10.1017/pan.2019.11. hdl:1721.1/128459.
  280. Fabijan, Aleksander; Gupchup, Jayant; Gupta, Somit; Omhover, Jeff; Qin, Wen; Vermeer, Lukas; Dmitriev, Pavel (2019-07-25). "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments". Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. pp. 2156–2164. doi:10.1145/3292500.3330722. ISBN 978-1-4503-6201-6.
  281. Kotok, Alan (19 March 2020). "WHO beginning Covid-19 therapy trial". Technology News: Science and Enterprise. Retrieved 7 April 2020.
  282. "Launch of a European clinical trial against COVID-19". INSERM. 22 March 2020. Retrieved 5 April 2020.
  283. Sert, Nathalie Percie du; Hurst, Viki; Ahluwalia, Amrita; Alam, Sabina; Avey, Marc T.; Baker, Monya; Browne, William J.; Clark, Alejandra; Cuthill, Innes C.; Dirnagl, Ulrich; Emerson, Michael (14 July 2020). "The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research". PLOS Biology. 18 (7) e3000410. doi:10.1371/journal.pbio.3000410. PMC 7360023. PMID 32663219.{{cite journal}}: CS1 maint: unflagged free DOI (link)
  284. O'Grady, Cathleen (14 July 2020). "Journals endorse new checklist to clean up sloppy animal research". Science. Retrieved 13 August 2026.
  285. Garcia-Pelegrin, Elias; Schnell, Alexandra K.; Wilkins, Clive; Clayton, Nicola S. (2020). "An unexpected audience". Science. 369 (6510): 1424–1426. doi:10.1126/science.abc6805. PMID 32943508.
  286. Adda, Jérôme; Decker, Christian; Ottaviani, Marco (2020-06-16). "P-hacking in clinical trials and how incentives shape the distribution of results across phases". Proceedings of the National Academy of Sciences of the United States of America. 117 (24): 13386–13392. arXiv:1907.00185. doi:10.1073/pnas.1919906117. PMID 32487730.
  287. Garcia-Pelegrin, Elias; Schnell, Alexandra K.; Wilkins, Clive; Clayton, Nicola S. (2021). "Exploring the perceptual inabilities of Eurasian jays (Garrulus glandarius) using magic effects". PNAS. 118 (24) e2026106118. doi:10.1073/pnas.2026106118. PMC 8214664. PMID 34074798.
  288. Busby, Mattha (27 June 2020). "Cambridge college to remove window commemorating eugenicist". The Guardian. Retrieved 2020-06-28.
  289. Erslev, Malthe Stavning (2022). "A Mimetic Method: Rendering Artificial Intelligence Imaginaries through Enactment". A Peer-Reviewed Journal About. 11 (1): 34–49. doi:10.7146/aprja.v11i1.134305.
  290. Clark, Cory J.; Costello, Thomas; Mitchell, Gregory; Tetlock, Philip E. (March 2022). "Keep your enemies close: Adversarial collaborations will improve behavioral science". Journal of Applied Research in Memory and Cognition. 11 (1): 1–18. doi:10.1037/mac0000004.
  291. "In the pursuit of scientific truth, working with adversaries can pay off". Penn Today. 7 July 2022. Retrieved 13 August 2026.
  292. Feng, Tianxing; Zhu, Chouwen; Yang, Geliang (2025). "Viewpoints of investigator on CONSORT 2025 statement-updated guideline for reporting randomized trials". Journal of Thoracic Disease. 17 (5): 2752–2755. doi:10.21037/jtd-2025-871. PMC 12170029. PMID 40529734. {{cite journal}}: Check |pmc= value (help); Check |pmid= value (help)CS1 maint: unflagged free DOI (link)
  293. Msaouel, Pavlos (2025-03-28). "The curious rise of randomised non-comparative trials". Significance. 22 (3): 40–44. doi:10.1093/jrssig/qmaf029. ISSN 1740-9705.
  294. Sherry, Alexander D.; Msaouel, Pavlos; Ludmir, Ethan B. (2024). "A meta-epidemiological analysis of post-hoc comparisons and primary endpoint interpretability among randomized noncomparative trials in clinical medicine". Journal of Clinical Epidemiology. 175 111540. doi:10.1016/j.jclinepi.2024.111540.
  295. "Image-based treatment effect heterogeneity". arXiv. Retrieved 3 April 2026.
  296. "Design of experiments". Google Trends. Retrieved 14 December 2021.
  297. "Design of experiments". books.google.com. Retrieved 14 December 2021.
  298. "Design of experiments". wikipediaviews.org. Retrieved 14 December 2021.