SymptomAI: Towards a conversational AI agent for everyday symptom assessment
Google Research introduces SymptomAI, a conversational AI agent based on Gemini Flash 2.0 for symptom assessment. In a randomized national study with 13,917 participants, five prompting strategies (dynamic, fixed canonical, flexible canonical, and unguided baseline) were compared against clinician assessments. Key findings: clinicians preferred SymptomAI's differential diagnosis (DDx) over peer clinicians' in >50% of cases; all agent-driven strategies significantly outperformed the unguided baseline in top-5 accuracy; AI's relative advantage was greatest for low-confidence clinician cases. SymptomAI diagnoses correlated with Fitbit biosignals (heart rate, skin temperature, sleep) around symptom onset, especially for respiratory infections. The work demonstrates AI's potential for real-world symptom assessment but stresses that all outputs are research-only.
A large proportion of clinical diagnoses can be derived from language-based interviews alone. These diagnostic interviews are typically conducted by clinicians through doctor-patient interactions during in-person or remote visits. While these interactions are the gold standard for symptom assessment, they can often suffer from financial, geographic, and systemic barriers that limit their accessibility. Current language models (LMs) have demonstrated strong differential diagnosis assessment capabilities when evaluated on curated medical case-studies, highlighting their potential to support the diagnostic process. However, existing evaluations have largely relied on curated, highly detailed and sometimes synthetic patient vignettes, which may not reflect real world experience and clinical presentation variability. These evaluations do not capture how everyday patients report their health symptoms, for example with varying levels of medical literacy, incomplete information, and other complexities that arise through natural conversation. This represents a key gap, leading to uncertainty of how LMs might perform in real-world contexts.
很大一部分临床诊断仅凭基于语言的访谈即可得出。这些诊断访谈通常由临床医生通过面对面或远程诊疗中的医患互动进行。尽管这些互动是症状评估的黄金标准,但往往受到经济、地域和系统性障碍的限制,从而影响了其可及性。当前的语言模型(LM)在基于精选医学案例研究进行评估时,已展现出强大的鉴别诊断能力,凸显了其支持诊断过程的潜力。然而,现有的评估大多依赖精心设计、高度详细甚至合成的患者场景,可能无法反映真实世界的经验和临床表现的多样性。这些评估未能捕捉到日常患者如何报告其健康症状,例如患者医疗素养参差不齐、信息不完整以及自然对话中出现的其他复杂性。这代表了一个关键差距,导致人们不确定LM在真实环境中表现如何。
To address this gap, we conduct an in-situ comparative research study of a set of experimental conversational prototype AI agents designed to explore how conversational AI might conduct end-to-end symptom interviews and differential diagnostic assessment for research benchmarking purposes. In our recent research paper, “SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment”, we share results from a randomized national scale study (n=13,917) in which consented research participants interact with one of five possible Gemini Flash 2.0 SymptomAI agents. All diagnoses, labels, and disease associations generated during the study were for research analysis only and did not constitute confirmed clinical diagnoses or official medical assessments. Two weeks after their interaction with the AI agents, we asked research participants to report any diagnoses they received from a visit with a healthcare provider. Using this data, we conducted a clinical expert annotation study comparing SymptomAI’s diagnostic performance relative to real clinicians' medical assessments.
为填补这一空白,我们开展了一项现场比较研究,探索一组实验性的对话式AI原型代理如何完成端到端的症状访谈和鉴别诊断评估,以供研究基准测试之用。在我们近期发表的论文《SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment》中,我们分享了一项全国范围随机研究(n=13,917)的结果。在该研究中,经过同意的研究参与者与五种Gemini Flash 2.0 SymptomAI代理之一进行交互。研究过程中生成的所有诊断、标签和疾病关联仅用于研究分析,不构成确诊的临床诊断或官方医学评估。在与AI代理互动两周后,我们要求研究参与者报告他们从医疗提供者处获得的任何诊断。利用这些数据,我们进行了一项临床专家标注研究,将SymptomAI的诊断表现与真实临床医生的医学评估进行比较。
After assessing the accuracy of SymptomAI’s differential diagnoses (DDx), we further compare SymptomAI’s diagnoses against biosignals from participants’ Fitbit wearable devices in the time leading up to their conversation with SymptomAI. We show that SymptomAI conversations that led to diagnosis with an infectious disease etiology coincide with physiological trends that may indicate an immune response, suggesting further evidence of SymptomAI’s performance.
在评估了SymptomAI鉴别诊断(DDx)的准确性之后,我们进一步将SymptomAI的诊断与参与者在与SymptomAI对话前来自Fitbit可穿戴设备的生物信号进行了比较。我们发现,导致传染病病因诊断的SymptomAI对话与可能表明免疫反应的生理趋势相吻合,这为SymptomAI的性能提供了进一步证据。
We enrolled 13,917 consenting research study participants who each describe their symptoms to one of five randomized SymptomAI agents, each with varying degrees of flexibility in how they conducted the symptom interview. During these conversations, participants described their symptoms and SymptomAI asked follow-up questions, with conversations culminating in a final differential diagnosis (DDx, a list of plausible diagnoses) and recommendations for next steps. Participants could then go on to see a healthcare provider and were asked to share the outcome of that visit via a survey two-weeks later. To evaluate and baseline SymptomAI’s assessment, we conducted a clinical-expert annotation study in which a panel of three board-certified clinicians reviewed the conversation transcripts and provided their own assessment (i.e., differential diagnosis). Then each clinician, in a blinded fashion, ranked the DDx provided by SymptomAI and those provided by the remaining clinicians.
我们招募了13,917名同意参与的研究参与者,每人向五种随机分配的SymptomAI代理之一描述其症状,这些代理在症状访谈方式上具有不同程度的灵活性。在对话过程中,参与者描述症状,SymptomAI提出后续问题,最终对话给出鉴别诊断(DDx,即一组可能的诊断)和下一步建议。参与者随后可以去看医疗提供者,并在两周后通过调查问卷告知就诊结果。为了评估SymptomAI的诊断并为临床诊断建立基线,我们进行了一项临床专家标注研究:由三名具备执业资格的临床医生组成评审小组,审查对话记录并提供他们自己的评估(即鉴别诊断)。然后,每位临床医生以盲法方式对SymptomAI提供的DDx和其他临床医生提供的DDx进行排名。
We found that the clinicians preferred the DDx generated by SymptomAI over those provided by the other clinicians in over 50% of the cases. This indicates that SymptomAI DDx aligned with our clinicians’ medical assessments just as often or more often than that of other clinicians.
我们发现,在超过50%的案例中,临床医生更偏好SymptomAI生成的DDx,而非其他临床医生提供的DDx。这表明SymptomAI的DDx与临床医生的医学评估的一致程度不亚于甚至高于其他临床医生。
Similarly, we compare the accuracy of the DDx generated by SymptomAI and provided by real clinicians via top-5 Accuracy (i.e., whether the true diagnosis provided by our participants' personal healthcare provider appears as one of the five possible diagnoses in the DDx). We had our clinicians each identify whether the provided diagnosis was in each DDx, including both the DDx generated by SymptomAI and those provided by clinicians. We found that the clinicians ranked the DDx generated by SymptomAI to be accurate more often than the DDx provided by other clinicians.
同样地,我们通过前五准确率(即参与者个人医疗提供者给出的真实诊断是否出现在DDx的五个可能诊断中)比较了SymptomAI生成的DDx和真实临床医生提供的DDx的准确性。我们让每位临床医生判断每个DDx中是否包含该真实诊断,包括SymptomAI生成的DDx和临床医生提供的DDx。结果发现,临床医生认为SymptomAI生成的DDx准确的频率高于其他临床医生提供的DDx。
As part of this research, we assessed different approaches for conducting history taking interviews. Participants were randomly assigned to five study arms, each employing different prompting strategies. Two (Dynamic Live and Dynamic Final) were given total agency to ask unrestricted follow up questions, two more (Fixed Canonical and Flexible Canonical) each asked questions from a set of standard history taking questions taught in medical school, and finally a Base unprompted LM, representing the fully user-driven experience that is the current status quo when querying LM chatbots. We found that all agent-driven prompting strategies (i.e., where SymptomAI actively asked follow up questions) significantly outperformed the Base condition, demonstrating the value of eliciting information from participants for improving differential diagnostic accuracy.
作为本研究的一部分,我们评估了不同的病史访谈方法。参与者被随机分配到五个研究组,每组采用不同的提示策略。其中两个组(Dynamic Live和Dynamic Final)被赋予完全自主权,可以问不受限制的后续问题;另外两个组(Fixed Canonical和Flexible Canonical)则从医学院教授的标准病史提问集中选择问题;最后一个为Base组,即无提示的LM,代表当前查询LM聊天机器人时完全由用户驱动的体验。我们发现,所有由代理驱动的提示策略(即SymptomAI主动提出后续问题)都显著优于Base条件,证明了主动从参与者处获取信息对于提高鉴别诊断准确性的价值。
We found that SymptomAI’s performance above clinical baselines was greatest for cases where the clinician’s felt least confident in their own DDx.
我们发现,在临床医生对自身DDx最没有信心的案例中,SymptomAI相对于临床基线的提升最为显著。
Given SymptomAI's accuracy against clinical baselines, we can also explore its potential at scale. Currently, the cost of clinical labels prohibits real-world analyses of population-scale datasets. Accurate symptom checking systems like SymptomAI have the potential to enable automated reference labeling of clinical quality diagnosis, which can open up large-scale analyses of physiological data — a task that is currently impossible at scale.
One such example is correlating wearable biosignals with different categories of illness. The most notable changes in wearable biosignals are observed for acute respiratory infections. To study this at population scale, we collected daily biometric data from our consenting participants for up to 30 days prior to their interaction with SymptomAI. We find clear biosignal shifts indicating symptom onset in the days approaching the user's symptom reporting. Importantly, the separation between cohorts was derived through categorizing SymptomAI's top-1 candidate diagnosis and grouping diagnoses that were classified as respiratory infections. This cohort excludes non-infectious respiratory illnesses like allergic rhinitis or chronic obstructive pulmonary disease. The correlation of wearable biosignals shift peaks aligning with the date of symptom reporting for these participants serves as observational physiological evidence that align with their reported symptoms.
鉴于SymptomAI相对于临床基线的准确性,我们还可以大规模探索其潜力。目前,临床标签的成本阻碍了对人口规模数据集的真实世界分析。准确的症状核查系统(如SymptomAI)有可能实现临床质量诊断的自动参考标注,从而开启大规模生理数据分析——这是一项目前无法大规模完成的任务。
一个例子是将可穿戴生物信号与不同疾病类别相关联。可穿戴生物信号最显著的变化见于急性呼吸道感染。为了在群体规模上研究这一点,我们收集了同意参与者在与SymptomAI互动前最多30天的每日生物特征数据。我们发现,在用户报告症状之前的几天里,生物信号出现了明显变化,表明症状发作。重要的是,队列之间的区分是通过分类SymptomAI排名第一的候选诊断并将被归类为呼吸道感染的诊断分组而得出的。该队列排除了非传染性呼吸道疾病,如过敏性鼻炎或慢性阻塞性肺疾病。可穿戴生物信号变化峰值与这些参与者报告症状的日期相一致,这为观察到的生理证据与报告的症状相符提供了支持。
AI-based assessment of symptom presentations opens the door to new research. By using SymptomAI to analyze a large volume of symptom reports and pairing those with real-time Fitbit data, we can explore digital biosignal phenotypes across a wide range of diseases. Our analysis revealed distinct shifts in physiological metrics — including cardiovascular function, respiration, skin temperature, and sleep quality — in the days leading up to a user's SymptomAI conversation. These objective changes align closely with the timing of the symptom conversation, offering a potential way to validate patient-reported symptoms or provide passive data to help inform a differential diagnosis alongside their symptom conversation. Additionally, this real-time accessibility highlights a core benefit of AI symptom checkers. Unlike traditional clinical appointments that can suffer from scheduling delays, participants could take part on the SymptomAI research study contemporaneously while symptoms are fresh. This potentially could improve the accuracy of patient-reported onset timelines — a crucial detail for population-scale health analysis.
基于AI的症状表现评估为新的研究打开了大门。通过使用SymptomAI分析大量症状报告,并将其与实时的Fitbit数据配对,我们可以探索跨多种疾病的数字生物信号表型。我们的分析揭示了用户在SymptomAI对话前几天的生理指标(包括心血管功能、呼吸、皮肤温度和睡眠质量)的明显变化。这些客观变化与症状对话的时间点紧密吻合,提供了一种潜在的途径来验证患者报告的症状,或提供被动数据以辅助鉴别诊断。此外,这种实时可及性凸显了AI症状检查器的核心优势。与可能因排期延迟的传统临床预约不同,参与者可以在症状出现时同步参与SymptomAI研究。这有可能提高患者报告发病时间线的准确性——这是群体规模健康分析的一个关键细节。
SymptomAI is an exploratory research effort that could represent a significant research advancement in AI-based symptom assessment and demonstrates the potential it could provide for the general public seeking understanding of their symptoms. While a population deployment evaluation reveals the accuracy of symptom assessment through remote patient interviews, there are nuanced limitations when comparing against clinician’s assessments.
Firstly, differential diagnosis itself is an ambiguous task and even reported diagnoses may change and develop longitudinally. A symptom assessment is a snapshot in time and captures the symptoms as they present in that moment. Due to the scale of our deployment, we were unable to control for frequency and timing of symptom reporting. As a result, some participants may have reported their symptoms well before more representative indicators developed, while others may have reported obvious indicators from an informed context after years of experience with chronic illness. Future work may focus on specific illnesses at specific points during symptom development such as early-onset metabolic syndrome or symptoms discussed at the start of respiratory infections. All diagnoses, labels, and disease associations generated during the study are AI-derived for research analysis only and do not constitute confirmed clinical diagnoses or official medical assessments.
Secondly, in our evaluation the clinicians reviewed static chat transcripts and were not given agency to ask their own follow-up questions. Clinicians may have intuitively sourced different information had they directed the symptom interview. Moreover, while recent research has shown that conversational AI systems can source clinical data with a clinician-level of detail and accuracy, such systems may miss alternative signals like body language, visual assessment, medical records, or in the context of primary care, existing rapport with the patient.
SymptomAI是一项探索性研究,可能代表了基于AI的症状评估领域的重要研究进展,并展示了其为寻求了解自身症状的公众提供的潜力。尽管人群部署评估显示了远程患者访谈在症状评估中的准确性,但在与临床医生的评估进行比较时仍存在微妙的局限性。
首先,鉴别诊断本身就是一个模糊的任务,甚至报告出的诊断也可能随时间推移而变化和发展。症状评估是时间点上的快照,捕捉的是该时刻呈现的症状。由于我们部署的规模,我们无法控制症状报告的频率和时机。因此,一些参与者可能在更典型指标出现之前就报告了症状,而另一些参与者则可能在多年慢性病经验后从知情背景中报告了明显指标。未来的工作可以聚焦于症状发展特定时间点的特定疾病,例如早期发作的代谢综合征或呼吸道感染初期讨论的症状。研究过程中生成的所有诊断、标签和疾病关联均为AI衍生,仅用于研究分析,不构成确诊的临床诊断或官方医学评估。
其次,在我们的评估中,临床医生审查的是静态聊天记录,没有被赋予提出自己后续问题的自主权。如果他们主导症状访谈,可能会直观地获取不同的信息。此外,尽管近期研究表明对话式AI系统能够以临床医生级别的细节和准确性获取临床数据,但此类系统可能遗漏其他信号,如肢体语言、视觉评估、医疗记录,或在初级保健环境中与患者已有的默契。
In conclusion, we introduce SymptomAI, an investigational conversational AI agent for conducting real-world patient interviews and symptom assessments. We demonstrate SymptomAI’s end-to-end real-world performance through DDx accuracy on a population sample, and show how SymptomAI diagnoses can enable analysis of population-scale signals like wearable biosignals for identifying associations in physiological signals with reported illness.
总之,我们介绍了SymptomAI,这是一个用于进行真实世界患者访谈和症状评估的研究性对话式AI代理。我们通过在人群样本上的DDx准确性展示了SymptomAI的端到端真实世界性能,并展示了SymptomAI诊断如何能够实现群体规模信号(如可穿戴生物信号)的分析,以识别生理信号与报告疾病之间的关联。