{"product_id":"artificial-intelligence-in-medicine-the-promise-the-pitfalls-and-what-patients-should-know","title":"Artificial Intelligence in Medicine: The Promise, the Pitfalls, and What Patients Should Know","description":"\u003cp\u003eIn a review published in the \u003cem\u003eNew England Journal of Medicine\u003c\/em\u003e, researchers from the University of Oxford examine the intersection of traditional medical statistics and artificial intelligence (AI), explaining how AI's greatest strength—automated pattern discovery in massive datasets—is also its greatest statistical vulnerability. The article explores key challenges including the gap between predicting individual outcomes and understanding population health, the difficulty of verifying AI's internal reasoning, and the risk that algorithms may inherit or amplify hidden biases, as demonstrated by a real-world algorithm applied to 200 million Americans that unintentionally discriminated against Black patients. The authors call for transparency, rigorous validation, and human oversight to ensure that AI in medicine is safe, reliable, and equitable.\u003c\/p\u003e\n\n\u003ch1\u003eArtificial Intelligence in Medicine: The Promise, the Pitfalls, and What Patients Should Know\u003c\/h1\u003e\n\n\u003ch2\u003eTable of Contents\u003c\/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"#ddn-key-points\"\u003eKey Points\u003c\/a\u003e\u003c\/li\u003e\n\n  \u003cli\u003e\u003ca href=\"#background\"\u003eWhy This Research Matters\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#history\"\u003eA Century of Statistics in Medicine\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#feature-learning\"\u003eHow AI Learns: The Power of Automated Pattern Detection\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#prediction-vs-inference\"\u003ePrediction vs. Understanding: Two Different Goals\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#generalizability\"\u003eWill AI Work in the Real World? Generalizability and Interpretation\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#stability\"\u003eKeeping AI Honest: Stability and Statistical Guarantees\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#bias\"\u003eWhen AI Gets It Wrong: Real-World Bias\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#clinical-implications\"\u003eWhat This Means for Patients\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#limitations\"\u003eLimitations of This Review\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#recommendations\"\u003eRecommendations for the Future\u003c\/a\u003e\u003c\/li\u003e\n\u003cli\u003e\u003ca href=\"#ddn-faq\"\u003eFrequently Asked Questions\u003c\/a\u003e\u003c\/li\u003e\n\u003cli\u003e\u003ca href=\"#source\"\u003eSource Information\u003c\/a\u003e\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003c!-- ddn:keypoints:start --\u003e\n\u003ch2 id=\"ddn-key-points\"\u003eKey Points\u003c\/h2\u003e\n\u003cul\u003e\n\u003cli\u003eAI is already used in medicine for mammograms, disease risk prediction, and electronic health records, but it has statistical vulnerabilities.\u003c\/li\u003e\n\u003cli\u003eAI's automated pattern discovery can amplify hidden biases; an algorithm applied to 200 million Americans discriminated against Black patients.\u003c\/li\u003e\n\u003cli\u003eAI models are often hard to interpret and verify, requiring transparency and human oversight for safe medical use.\u003c\/li\u003e\n\u003cli\u003eThere is a gap between AI's ability to predict individual outcomes and understanding disease in the broader population.\u003c\/li\u003e\n\u003cli\u003eReleasing code, prespecified analysis plans, combining AI with conventional statistics, and human judgment are recommended for responsible AI use.\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003c!-- ddn:keypoints:end --\u003e\n\n\n\u003ch2 id=\"background\"\u003eWhy This Research Matters\u003c\/h2\u003e\n\n\u003cp\u003eArtificial intelligence is no longer a distant concept in healthcare—it is already here. AI systems are being used to read mammograms, predict disease risk, analyze electronic health records, and even assist with medical note-taking. But how do we know these systems are safe, accurate, and fair?\u003c\/p\u003e\n\n\u003cp\u003eThis review article, written by David J. Hunter, M.B., B.S., from the Nuffield Department of Population Health at the University of Oxford, and Christopher Holmes, Ph.D., from the Department of Statistics and the Alan Turing Institute in London, tackles a critical question: \u003cstrong\u003ewhat happens when the rigorous world of medical statistics meets the powerful but often opaque world of artificial intelligence?\u003c\/strong\u003e\u003c\/p\u003e\n\n\u003cp\u003eThe authors describe a fundamental paradox. AI's ability to automatically discover patterns in enormous datasets makes it an incredibly valuable medical research tool—but the very same features make it statistically vulnerable. Techniques adequate for targeted advertising to voters or consumers may not meet the rigorous demands of risk prediction or diagnosis in medicine. This article explores that tension and what it means for the future of patient care.\u003c\/p\u003e\n\n\u003ch2 id=\"history\"\u003eA Century of Statistics in Medicine\u003c\/h2\u003e\n\n\u003cp\u003eStatistics emerged as a distinct discipline around the beginning of the \u003cstrong\u003e20th century\u003c\/strong\u003e. During this period, fundamental concepts were developed that would transform medical research, including the use of randomization in clinical trials, hypothesis testing, likelihood-based inference, P values, and Bayesian analysis and decision theory.\u003c\/p\u003e\n\n\u003cp\u003eStatistics quickly became essential to applied sciences. In fact, in \u003cstrong\u003e2000, the editors of the \u003cem\u003eNew England Journal of Medicine\u003c\/em\u003e cited \"Application of Statistics to Medicine\" as one of the 11 most important developments in medical science over the previous 1000 years.\u003c\/strong\u003e\u003c\/p\u003e\n\n\u003cp\u003eSo what exactly is statistics? The authors define it as reasoning with incomplete information—the rigorous interpretation and communication of scientific findings from data. Statistics includes determining the optimal design of experiments and accurately quantifying uncertainty about conclusions, all expressed through the language of probability.\u003c\/p\u003e\n\n\u003cp\u003eNow, in the 21st century, artificial intelligence has emerged as a powerful new force in medical research. This development is driven, in part, by enormous expansions in computer power and data availability—but with these advances come new statistical challenges that this review sets out to address.\u003c\/p\u003e\n\n\u003ch2 id=\"feature-learning\"\u003eHow AI Learns: The Power of Automated Pattern Detection\u003c\/h2\u003e\n\n\u003ch3\u003eThe Traditional Approach: Hands-On Statistics\u003c\/h3\u003e\n\n\u003cp\u003eTraditional statistical modeling relies on careful, hands-on selection of measurements and data features to include in an analysis. For example, a statistician must decide which covariates (factors that might influence the outcome) to include in a regression model—a statistical method that examines relationships between variables. They also determine what transformations or standardizations of measurements are needed.\u003c\/p\u003e\n\n\u003cp\u003eSemiautomated data-reduction techniques such as \u003cstrong\u003erandom forests\u003c\/strong\u003e and forward- or backward-selection stepwise regression have assisted statisticians in this selection process for decades. In this traditional framework, modeling assumptions and features are typically explicit, and the number of parameters in the model is usually known.\u003c\/p\u003e\n\n\u003ch3\u003eThe AI Revolution: Automatic Feature Learning\u003c\/h3\u003e\n\n\u003cp\u003eArguably the most impressive and distinguishing aspect of AI is its automated ability to search and extract arbitrary, complex, task-oriented features from data—a process called \u003cstrong\u003efeature representation learning\u003c\/strong\u003e. Features are algorithmically engineered from data during a training phase to uncover data transformations that are correct for the learning task.\u003c\/p\u003e\n\n\u003cp\u003eAI algorithms largely remove the need for analysts to prespecify features or manually curate variable transformations. This is especially beneficial in large, complex data domains such as:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eImage analysis (like mammograms or pathology slides)\u003c\/li\u003e\n  \u003cli\u003eGenomics (studying the complete set of genes in an organism)\u003c\/li\u003e\n  \u003cli\u003eModeling electronic health records (digital versions of patients' medical histories)\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003cp\u003eAI models can search through potentially billions of nonlinear covariate transformations to reduce a large number of variables to a smaller set of task-adapted features. Somewhat paradoxically, increasing the complexity of an AI model through additional parameters—which is what happens in deep learning—actually helps the model find richer internal feature sets, provided training methods are suitably tailored.\u003c\/p\u003e\n\n\u003ch3\u003eThe Dark Side of the Paradox\u003c\/h3\u003e\n\n\u003cp\u003eHere is the problem: the features AI engineers are often \u003cstrong\u003ebeyond the scope of what humans can create\u003c\/strong\u003e, which is why AI performs so impressively. But those same features are often hard to interpret, are brittle (prone to breaking) when data changes, and lack common sense when it comes to using background knowledge and qualitative checks that statisticians routinely apply.\u003c\/p\u003e\n\n\u003cp\u003eAI models are often unable to trace the evidence line from data to features, making auditability and verification challenging. This means greater checks and balances are needed to ensure the validity and generalizability of AI-enabled scientific findings.\u003c\/p\u003e\n\n\u003cp\u003eThe authors also briefly comment on the emerging field of \u003cstrong\u003egenerative AI\u003c\/strong\u003e—such as large language models and medical chatbots that might be used for medical note-taking in electronic health records. These \"foundation models\" use self-supervised learning on vast quantities of undocumented training data, with objective functions trained using \u003cstrong\u003etrillions of parameters\u003c\/strong\u003e (at the time of writing). This is a stark contrast to \"supervised\" learning, where training data are known and labeled according to clinical outcomes, and the training objective is clear and targeted. Given the opaqueness of these generative models, the authors urge extra caution for health applications.\u003c\/p\u003e\n\n\u003ch2 id=\"prediction-vs-inference\"\u003ePrediction vs. Understanding: Two Different Goals\u003c\/h2\u003e\n\n\u003cp\u003eAI is especially well suited to—and largely designed for—\u003cstrong\u003elarge-scale prediction tasks\u003c\/strong\u003e. The training objective is clear, and predictive accuracy is usually well characterized. A good example is predicting the risk of disease.\u003c\/p\u003e\n\n\u003cp\u003eHowever, the ultimate goal of most medical studies is \u003cem\u003enot\u003c\/em\u003e explicitly to predict risk. Rather, it is to understand some biological mechanism or cause of disease in the wider population, or to assist in developing new therapies. The authors emphasize that there is an \u003cstrong\u003eevidence gap between a good predictive model that operates at the individual level and the ability to make inferential statements about the population\u003c\/strong\u003e.\u003c\/p\u003e\n\n\u003cp\u003eStatistics is mainly concerned with population inference—generalizing evidence from one study to a scientific hypothesis about the broader population. Prediction is an important yet simpler task; scientific inference often has a greater influence on mechanistic understanding.\u003c\/p\u003e\n\n\u003cp\u003eAs Hippocrates observed centuries ago: \u003cem\u003e\"It is more important to know what sort of person has a disease than to know what sort of disease a person has.\"\u003c\/em\u003e\u003c\/p\u003e\n\n\u003ch3\u003eA Real-World Example: The Covid-19 Pandemic\u003c\/h3\u003e\n\n\u003cp\u003eDuring the Covid-19 pandemic, various prediction tools were developed to determine whether a person had SARS-CoV-2 (the virus that causes Covid-19) infection. However, moving from individual prediction to understanding the population prevalence—and identifying which subgroups in the population were at higher risk—proved much more challenging.\u003c\/p\u003e\n\n\u003ch3\u003eHow Do We Measure Predictive Accuracy?\u003c\/h3\u003e\n\n\u003cp\u003eAn additional challenge is that there are \u003cstrong\u003emany ways to measure and report predictive accuracy\u003c\/strong\u003e, including:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eArea under the receiver-operating-characteristic curve (a measure of how well a test distinguishes between groups)\u003c\/li\u003e\n  \u003cli\u003ePrecision and recall (measures of exactness and completeness)\u003c\/li\u003e\n  \u003cli\u003eMean squared error (average of the squares of the errors)\u003c\/li\u003e\n  \u003cli\u003ePositive predictive value (probability that a positive test result is correct)\u003c\/li\u003e\n  \u003cli\u003eMisclassification rate (how often the model gets the answer wrong)\u003c\/li\u003e\n  \u003cli\u003eNet reclassification index (whether a new model improves classification)\u003c\/li\u003e\n  \u003cli\u003eLog probability score (a measure of prediction confidence)\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003cp\u003eChoosing the measure that is appropriate for the context is vitally important, because accuracy in one measure may not translate to accuracy in another—and may not relate to a clinically meaningful measure of performance or safety.\u003c\/p\u003e\n\n\u003cp\u003eIn contrast, inferential targets for population statistics tend to be less ambiguous, and uncertainty is clearly characterized using \u003cstrong\u003eP values, confidence intervals, and credible intervals\u003c\/strong\u003e. Even so, robust, accurate AI prediction models indicate that repeatable signals and stable associations in the data are worth investigating further.\u003c\/p\u003e\n\n\u003ch3\u003eCausal Machine Learning: Where AI and Statistics Meet\u003c\/h3\u003e\n\n\u003cp\u003eAn interesting meeting point between AI prediction methods and statistical inference is \u003cstrong\u003ecausal machine learning\u003c\/strong\u003e, which pays particular attention to inferential quantities. By adopting structural causal modeling or potential outcomes frameworks—with tools such as directed acyclic graphs (diagrams that show cause-and-effect relationships)—researchers can use domain knowledge to reduce the probability that an AI model will make mistakes. These mistakes include misspecifying the temporal relationship between exposure and outcome, conditioning on a variable that is caused by both exposure and disease (a \"collider\"), or highlighting spurious associations (like a batch effect in a biomarker study).\u003c\/p\u003e\n\n\u003cp\u003eCausal inference methods may also be applied to AI to help interpret radiological or pathological images and to support clinical decision making and diagnosis. However, the authors stress that \u003cstrong\u003ehuman judgment will likely be necessary for the foreseeable future\u003c\/strong\u003e, if only because different AI algorithms may present us with different conclusions. Moreover, causal analysis from observational data requires assumptions that lie outside what can be learned from the data alone—to avoid bias from ascertainment, mediation, and confounding.\u003c\/p\u003e\n\n\u003ch2 id=\"generalizability\"\u003eWill AI Work in the Real World? Generalizability and Interpretation\u003c\/h2\u003e\n\n\u003cp\u003eOne of the biggest challenges in interpreting AI results is that algorithms for internal feature representation are designed to automatically adapt their complexity to the task at hand, with nearly infinite flexibility in some approaches.\u003c\/p\u003e\n\n\u003cp\u003eThis flexibility is a great strength—but it also requires care to avoid \u003cstrong\u003eoverfitting\u003c\/strong\u003e, which happens when a model learns the noise and random quirks of its training data so well that it performs poorly on new, unseen data.\u003c\/p\u003e\n\n\u003ch3\u003eThe Limits of Traditional Statistical Guarantees\u003c\/h3\u003e\n\n\u003cp\u003eThe use of regularization (techniques that prevent overfitting by penalizing model complexity) and controlled stochastic optimization of model parameters during training can help prevent overfitting. But these techniques also mean AI algorithms have poorly defined notions of \u003cstrong\u003estatistical degrees of freedom\u003c\/strong\u003e and the number of free parameters. Traditional statistical guarantees against overoptimism cannot be used in the AI context.\u003c\/p\u003e\n\n\u003cp\u003eInstead, researchers must substitute techniques such as cross-validation and held-out samples to mimic true out-of-sample performance. The trade-off is that the amount of data available for discovery is reduced. Taken together, these factors create a real risk of \u003cstrong\u003eoverinterpreting the generalizability and reproducibility of results\u003c\/strong\u003e.\u003c\/p\u003e\n\n\u003ch3\u003eA Notable Example: AI and Breast Cancer Screening\u003c\/h3\u003e\n\n\u003cp\u003eThe authors highlight a well-known case involving a study by McKinney et al. on using AI to predict breast cancer based on mammograms. The study showed high potential for AI in breast cancer screening. However, Haibe-Kains and colleagues publicly called for greater transparency, stating: \u003cem\u003e\"In their study, McKinney et al. showed the high potential of AI for breast cancer screening. However, the lack of details of the methods and algorithm code undermines its scientific value.\"\u003c\/em\u003e\u003c\/p\u003e\n\n\u003cp\u003eThis illustrates why clear reporting of results and availability of code are essential for external replication and refinement by other research groups—though this may be limited by a tendency to seek intellectual property rights for commercial AI products.\u003c\/p\u003e\n\n\u003ch3\u003eUsing AI to Winnow Down Big Data\u003c\/h3\u003e\n\n\u003cp\u003eAI approaches can be helpful in reducing a dataset with a very large number of features—such as \"-omic\" datasets (metabolomic, proteomic, or genomic)—into a smaller number of features that can then be tested with conventional statistical methods.\u003c\/p\u003e\n\n\u003cp\u003ePopular AI methods that provide \"feature relevance\" rankings of covariates include:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cstrong\u003eRandom forests\u003c\/strong\u003e (an ensemble of decision trees)\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eXGBoost\u003c\/strong\u003e (a gradient-boosting algorithm)\u003c\/li\u003e\n  \u003cli\u003e\u003cstrong\u003eBayesian additive regression trees\u003c\/strong\u003e\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003cp\u003eStatistical methods such as the \u003cstrong\u003eleast absolute shrinkage and selection operator (LASSO)\u003c\/strong\u003e use explicit variable selection as part of model fitting. Feature reduction helps the human analyst examine the data more effectively and apply constraints based on previous subject knowledge. For example, a researcher might know that feature A is often confounded by feature X, or that a latency period of several years between exposure to feature A and the disease outcome means no relationship is expected in early follow-up.\u003c\/p\u003e\n\n\u003ch3\u003eAI vs. Conventional Statistics: A Side-by-Side Comparison\u003c\/h3\u003e\n\n\u003cp\u003eThe article provides a detailed comparison of AI methods and conventional statistics. Key differences include:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cstrong\u003ePrior hypotheses:\u003c\/strong\u003e AI is agnostic or very general about hypotheses; conventional statistics specifies hypotheses as primary, secondary, or exploratory.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eTechniques:\u003c\/strong\u003e AI uses random forests, neural networks, and XGBoost; conventional statistics uses parametric and nonparametric comparisons, regression, and survival models with linear predictors.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eStability:\u003c\/strong\u003e AI analyses are more prone to instability due to application domains and user choices in algorithm specification; conventional statistics follow a prespecified analysis plan with minimal user-defined choices.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eApplications:\u003c\/strong\u003e AI excels at images, monitor outputs, electronic health records, and natural language processing; conventional statistics suits data with fewer predictors, tabular data, and randomized trials.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003ePurpose:\u003c\/strong\u003e AI focuses on pattern discovery, automatic feature representation, feature reduction, and prediction; conventional statistics focuses on inference, testing specific factors, controlling confounding, and quantifying uncertainty.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eReproducibility:\u003c\/strong\u003e AI often provides internal reproducibility (cross-validation or split samples); conventional statistics ideally provides external reproducibility with new data.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eBarriers:\u003c\/strong\u003e AI faces proprietary algorithms not available to other researchers and unclear reporting; conventional statistics faces slow progress in sharing primary data.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eInterpretability:\u003c\/strong\u003e AI is often a black box; conventional statistics has explicit features and clear degrees of freedom.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eEquity:\u003c\/strong\u003e AI's data-driven feature learning is susceptible to biases in data, compounding health inequities; conventional statistics models are more easily checked for equity when relevant data are available.\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003ch2 id=\"stability\"\u003eKeeping AI Honest: Stability and Statistical Guarantees\u003c\/h2\u003e\n\n\u003cp\u003eMedical science is an iterative process of observation and hypothesis refinement—cycles of experimentation, analysis, and conjecture that lead to further experiments and ultimately to a level of evidence that refutes existing theories or supports new therapies and lifestyle recommendations.\u003c\/p\u003e\n\n\u003cp\u003eRandomized trials of investigational drugs have historically been held to a high standard of rigor. Concerns about overinterpretation of secondary end-point and subgroup analyses have led to an even stronger focus on prespecified description of primary hypotheses and control of the \u003cstrong\u003efamilywise error rate\u003c\/strong\u003e (the probability of making at least one false positive finding among multiple tests) to limit false positive results.\u003c\/p\u003e\n\n\u003cp\u003eProtocols now often specify:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eThe precise estimands (the exact quantities being measured)\u003c\/li\u003e\n  \u003cli\u003eThe methods of analysis that will be used to obtain P values\u003c\/li\u003e\n  \u003cli\u003eThe covariates to be controlled for\u003c\/li\u003e\n  \u003cli\u003eEven the dummy tables that will be filled in once data are complete\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003cp\u003eAnalyses in observational studies are usually less rigorously prespecified, although a \u003cstrong\u003estatistical analysis plan\u003c\/strong\u003e established before data analysis begins is increasingly expected as supplementary material in published reports.\u003c\/p\u003e\n\n\u003ch3\u003eThe AI Problem: Patterns Without Prespecification\u003c\/h3\u003e\n\n\u003cp\u003eAI approaches often seek patterns in data that are \u003cem\u003enot\u003c\/em\u003e prespecified—which is one of their strengths. But the downside is that the potential for false positive results increases unless rigorous procedures for assessing reproducibility are incorporated.\u003c\/p\u003e\n\n\u003cp\u003eNew reporting guidelines and recommendations for AI in medical science have been established to ensure greater trust and generalizability of conclusions. The authors also note that highly adaptive AI algorithms \u003cstrong\u003einherit all the biases and unrepresentativeness that might be present in the training data\u003c\/strong\u003e. When using black-box AI prediction tools, it can be difficult to judge whether predictive signals arise from confounding due to hidden biases in the data.\u003c\/p\u003e\n\n\u003cp\u003eMethods from the field of \u003cstrong\u003eexplainable AI (XAI)\u003c\/strong\u003e can help counter opaque feature representation learning. However, for applications in which safety is a critical issue, the black-box nature of AI models warrants careful consideration and justification.\u003c\/p\u003e\n\n\u003ch2 id=\"bias\"\u003eWhen AI Gets It Wrong: Real-World Bias\u003c\/h2\u003e\n\n\u003cp\u003eOne of the most striking examples of AI bias in healthcare comes from research by Obermeyer and colleagues. The team described an AI-informed algorithm that was applied to a population of \u003cstrong\u003e200 million persons in the United States each year\u003c\/strong\u003e to identify patients at the highest risk for incurring substantial health care costs and to refer them to \"high-risk care management programs.\"\u003c\/p\u003e\n\n\u003cp\u003eThe analysis suggested that the algorithm \u003cstrong\u003eunintentionally discriminated against Black patients\u003c\/strong\u003e. The reason? At every level of health care expenditure and age, \u003cstrong\u003eBlack patients had more coexisting conditions\u003c\/strong\u003e than White patients. Because the algorithm was trained on health care spending as a proxy for health needs—and because systemic factors lead to lower spending on Black patients even when they are sicker—the algorithm systematically underestimated the health needs of Black patients.\u003c\/p\u003e\n\n\u003cp\u003eThis example dramatically illustrates a core concern the authors raise: AI systems trained on real-world data can perpetuate and even amplify existing inequities in health care.\u003c\/p\u003e\n\n\u003ch2 id=\"clinical-implications\"\u003eWhat This Means for Patients\u003c\/h2\u003e\n\n\u003cp\u003eFor patients, the message is both hopeful and cautionary. AI has the potential to transform medicine in powerful ways:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eReading and interpreting medical images like mammograms with high accuracy\u003c\/li\u003e\n  \u003cli\u003ePredicting disease risk from electronic health records\u003c\/li\u003e\n  \u003cli\u003eAnalyzing genomic and other \"-omic\" data to uncover disease mechanisms\u003c\/li\u003e\n  \u003cli\u003eReducing massive datasets to the most relevant features for clinical questions\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003cp\u003eBut patients should also be aware of the limitations:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eAI results can be difficult to interpret and verify\u003c\/li\u003e\n  \u003cli\u003eAI models may not generalize to different populations than what they were trained on\u003c\/li\u003e\n  \u003cli\u003eBiases in training data can lead to unfair or inaccurate results for certain groups\u003c\/li\u003e\n  \u003cli\u003eThe lack of transparency in commercial algorithms makes independent validation harder\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003cp\u003ePatients can advocate for themselves by asking questions about how AI is being used in their care. Willingness by health care institutions to provide clear answers about the validation, limitations, and oversight of AI tools is a good sign of responsible implementation.\u003c\/p\u003e\n\n\u003cp\u003eThe authors note that using conventional statistical prediction methods alongside interpretable AI methods can help provide an understanding of prediction signals and can mitigate nonsensical associations. Combining the best of both approaches—AI's power to handle vast data and statistics' rigor in interpretation—is likely the safest path forward.\u003c\/p\u003e\n\n\u003ch2 id=\"limitations\"\u003eLimitations of This Review\u003c\/h2\u003e\n\n\u003cp\u003eAs a review article rather than an original research study, this paper synthesizes existing knowledge rather than presenting new experimental data. The authors acknowledge that space constraints precluded a detailed discussion of the important area of AI and experimental design, and they only briefly comment on the emerging area of generative AI and medical chatbots rather than providing a deep dive.\u003c\/p\u003e\n\n\u003cp\u003eThe article reflects the perspectives of two researchers based in the United Kingdom—one from population health at Oxford and one from statistics and medicine at Oxford and the Alan Turing Institute—and draws heavily on examples from the U.S. and U.K. health care contexts. The field of AI in medicine is evolving rapidly, and some specifics (such as the scale of foundation models with trillions of parameters) may change quickly over time.\u003c\/p\u003e\n\n\u003ch2 id=\"recommendations\"\u003eRecommendations for the Future\u003c\/h2\u003e\n\n\u003cp\u003eBased on their analysis, the authors offer several practical recommendations for medical scientists and health care institutions:\u003c\/p\u003e\n\n\u003col\u003e\n  \u003cli\u003e\n\u003cstrong\u003eRelease all code.\u003c\/strong\u003e Sharing code and providing clear statements on model fitting and held-out data used for reporting accuracy facilitates external assessment of reproducibility.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eBe transparent about data.\u003c\/strong\u003e Clear reporting of results and availability of code add to the potential for external replication and refinement by other groups.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eUse prespecified analysis plans.\u003c\/strong\u003e For observational studies, a statistical analysis plan established before data analysis should be expected, just as it is for randomized trials.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eCombine AI with conventional methods.\u003c\/strong\u003e Using traditional statistical prediction methods alongside interpretable AI methods can clarify prediction signals and mitigate nonsensical associations.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eEmploy feature reduction wisely.\u003c\/strong\u003e AI can reduce high-dimensional data to a smaller set of features that can then be tested with conventional statistics, allowing human analysts to apply subject-specific knowledge.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eMaintain human oversight.\u003c\/strong\u003e Human judgment will be necessary for the foreseeable future, especially because different AI algorithms may present different conclusions, and causal inference requires assumptions beyond what the data alone can reveal.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eFollow reporting guidelines.\u003c\/strong\u003e New reporting guidelines and recommendations for AI in medical science are being established to ensure greater trust and generalizability of conclusions.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eAssess internal reproducibility.\u003c\/strong\u003e Data should be partitioned into discovery and test sets, and generalizability to other datasets must be carefully evaluated.\u003c\/li\u003e\n\u003c\/ol\u003e\n\n\u003cp\u003eThe authors emphasize that approaching AI in medicine with appropriate rigor and safeguards is not just a technical concern—it is essential for patient safety, scientific integrity, and health equity.\u003c\/p\u003e\n\n\u003cp\u003eAs the article's framing makes clear, a technique that works for targeted advertising or weather prediction may not meet the rigorous demands of medical diagnosis and risk prediction. The stakes are simply too high.\u003c\/p\u003e\n\n\u003c!-- ddn:faq:start --\u003e\n\u003ch2 id=\"ddn-faq\"\u003eFrequently Asked Questions\u003c\/h2\u003e\n\u003ch3\u003eWhat is artificial intelligence (AI) currently used for in medicine?\u003c\/h3\u003e\n\u003cp\u003eAI is already used in healthcare to read mammograms, predict disease risk, analyze electronic health records, and assist with medical note-taking. It can also help analyze genomic data and reduce massive datasets to the most relevant features for clinical questions. These applications show AI's potential, but they come with important limitations to understand.\u003c\/p\u003e\n\u003ch3\u003eHow is AI different from traditional statistics in medical research?\u003c\/h3\u003e\n\u003cp\u003eTraditional statistics relies on careful, hands-on selection of measurements and features, with explicit modeling assumptions. AI automatically searches for patterns in vast datasets, learning features on its own. This makes AI powerful for prediction but can make it hard to interpret, verify, and check for hidden biases compared to conventional statistical methods.\u003c\/p\u003e\n\u003ch3\u003eCan AI be biased or make mistakes in healthcare?\u003c\/h3\u003e\n\u003cp\u003eYes. A real-world AI algorithm applied to 200 million Americans to identify high-risk patients unintentionally discriminated against Black patients. Because it was trained on healthcare spending, and systemic factors lead to lower spending on Black patients even when sicker, it underestimated their health needs. AI can inherit and amplify existing biases in training data.\u003c\/p\u003e\n\u003ch3\u003eWhy is it hard to verify AI's decisions in medicine?\u003c\/h3\u003e\n\u003cp\u003eAI models often use internal feature representation learning, which creates features that are hard to interpret and lack common sense. They are often unable to trace the evidence line from data to features, making auditability and verification challenging. This is why transparency, clear reporting, and human oversight are essential when AI is used in patient care.\u003c\/p\u003e\n\u003ch3\u003eWhat should patients ask about AI used in their care?\u003c\/h3\u003e\n\u003cp\u003ePatients can ask how AI is being used in their care, how it was validated, what its limitations are, and who oversees it. A health institution that gives clear answers about these topics is showing responsible implementation. Patients should also be aware that AI results may not generalize to different populations and can contain biases.\u003c\/p\u003e\n\u003ch3\u003eCan AI be used to understand disease causes, not just predict risk?\u003c\/h3\u003e\n\u003cp\u003eAI is especially good at large-scale prediction, like predicting disease risk. However, understanding disease mechanisms in the wider population is harder. There is an evidence gap between predicting outcomes for an individual and making inferential statements about the population. Combining AI with conventional statistics and human judgment is the safest path forward.\u003c\/p\u003e\n\u003ch3\u003eWhat does the future of AI in medicine look like for patients?\u003c\/h3\u003e\n\u003cp\u003eThe future holds both promise and caution. AI has potential to transform medicine, but patients should be aware of limitations like difficult interpretation, poor generalizability, and bias. Responsible implementation includes releasing code, transparency about data, using prespecified analysis plans, and human oversight. Combination of AI with conventional methods is likely the safest approach.\u003c\/p\u003e\n\u003ch3\u003eShould I get a second opinion if AI is used in my medical diagnosis or treatment plan?\u003c\/h3\u003e\n\u003cp\u003eYes, a second opinion can be valuable when AI tools are involved in your care. AI systems can be biased, as shown by an algorithm that unintentionally discriminated against Black patients, and their results can be hard to interpret or verify. A second opinion from an independent expert can help ensure your diagnosis or treatment plan is accurate and appropriate, especially if you have concerns about how AI was used. Diagnostic Detectives Network provides independent expert second opinions.\u003c\/p\u003e\n\u003c!-- ddn:faq:end --\u003e\n\n\u003ch2 id=\"source\"\u003eSource Information\u003c\/h2\u003e\n\n\u003cp\u003e\u003cstrong\u003eOriginal article title:\u003c\/strong\u003e Where Medical Statistics Meets Artificial Intelligence\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors:\u003c\/strong\u003e David J. Hunter, M.B., B.S., and Christopher Holmes, Ph.D.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eJournal:\u003c\/strong\u003e \u003cem\u003eThe New England Journal of Medicine\u003c\/em\u003e, 2023; volume 389, pages 1211–1219\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003ePublication date:\u003c\/strong\u003e September 28, 2023\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eDOI:\u003c\/strong\u003e 10.1056\/NEJMra2212850\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eArticle type:\u003c\/strong\u003e Review Article in the \"AI in Medicine\" series, edited by Jeffrey M. Drazen, M.D., with guest editors Isaac S. Kohane, M.D., Ph.D., and Tze-Yun Leong, Ph.D.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eCopyright:\u003c\/strong\u003e © 2023 Massachusetts Medical Society. All rights reserved.\u003c\/p\u003e\n\u003cp\u003e\u003cem\u003eThis patient-friendly article is based on peer-reviewed research and is intended for educational purposes. It does not constitute medical advice.\u003c\/em\u003e\u003c\/p\u003e","brand":"DiagnosticDetectives.Com","offers":[{"title":"Default Title","offer_id":47459220193436,"sku":null,"price":0.0,"currency_code":"DKK","in_stock":true}],"url":"https:\/\/diagnosticdetectives.dk\/products\/artificial-intelligence-in-medicine-the-promise-the-pitfalls-and-what-patients-should-know","provider":"DiagnosticDetectives.Com","version":"1.0","type":"link"}