This is an old revision of the document!
AI models in healthcare have shown strong predictive performance, which has not translated into better care or lower cost — a gap often called the AI chasm. The value of using an AI model is determined by the action pairing it enables: the lead time of the model's prediction or recommendation, whether an effective intervention exists, the work capacity to deliver it, and whether the resulting allocation of resources is fair. The focus, then, is not about building a better model. It is delivery science: making models workflow-aware, evaluating usefulness and fairness during development rather than after deployment, and bringing AI to the clinic safely, ethically, and cost-effectively. There are three efforts aimed at this goal.
This is the academic lab, which is a mix of doctors, engineers, informatics professionals, and students developing methods to learn from patient-level health data, answer clinical questions at the point of care, and research safe, ethical, and cost-effective use of predictive models. It is affiliated with the Department of Medicine, the Clinical Excellence Research Center, and the Department of Biomedical Data Science. The current research is organized around five themes, anchored on live systems and pairing research methodology with a real deployment to study: (1) evaluation infrastructure for clinical AI, including MedHELM and HealthAdminBench; (2) safety, reliability, and guardrails for generative clinical AI; (3) post-deployment learning from ChatEHR; (4) EHR foundation models and long-context representation; (5) rare disease phenotyping and discovery. Prior work at http://shahlab.stanford.edu/examples_of_prior_work and http://shahlab.stanford.edu/rail
The GUIDE-AI group, which stands for Guidance for the Use, Implementation, Development, and Evaluation of Healthcare AI. Its three-point mission is to build and deploy healthcare AI solutions, establish evaluation and monitoring methods for deployed AI tools, and disseminate actionable learnings to support high-value AI-augmented care. This group is where methods meet the health system's policy and oversight needs. The work spans target product profiles that specify the performance an AI tool must reach to produce benefit in a given setting, pragmatic and risk-proportionate evaluation methodology for predictive, generative, and agentic tools, and the under-examined downstream effects of AI use such as financial toxicity when AI-recommended testing precedes payer policy. Resources at http://guide-ai.stanford.edu/resources, papers at http://guide-ai.stanford.edu/our-work.
The data science team, is an interdisciplinary team in Technology and Digital Solutions (TDS) started in March 2022, focused on ensuring Stanford Health Care is a leader in responsible AI; from development and implementation to maintenance and optimization. The work spans identification of use cases, deployment, and the MLOps that determine whether a model helps in routine clinical and operational use. Most recent effort is ChatEHR (http://chatehr.stanford.edu), which brings a conversational interaction to a medical record. Four areas of the group’s activity are at tds.stanfordmedicine.org/datascience