Level C· Early human research exploring benefitsProspective StudyEurope PMCOpen access

Comparative Performance of Current Patient-Accessible Artificial Intelligence Large Language Models in the Preoperative Education of Patients in Facial Aesthetic Surgery

Abi-Rafeh J., Bassiri-Tehrani B., Kazan R., Hanna SA., Kanevsky J., Nahai F.

Prospective Study on Face & Skin, published in Aesthet Surg J Open Forum (2024) — summary generated from the PubMed abstract.

Open my reading list
Level C· Early human research exploring benefitsEvidence level of this study

Early human evidence such as case series or small samples is exploring possible benefits.

  • Level A · Stronger Clinical Evidence
  • Level B · Emerging clinical evidence with positive signals
  • Level C · Early human research exploring benefits
  • Level D · Scientific groundwork from lab and animal studies
  • Emerging · Emerging topic under active research
Read the A–D evidence level guide

This page is generated from the PubMed record. The Thai description is an automated summary of bibliographic fields and the abstract, not a full translation, and is not medical advice.

Study type
Prospective Study
Journal
Aesthet Surg J Open Forum (2024)
Reported sample size
—
Source database
Europe PMC
PMID
39228821
PMCID
PMC11371156
DOI
10.1093/asjof/ojae058
Citations
5

Abstract (original English)

Background Artificial intelligence large language models (LLMs) represent promising resources for patient guidance and education in aesthetic surgery. Objectives The present study directly compares the performance of OpenAI's ChatGPT (San Francisco, CA) with Google's Bard (Mountain View, CA) in this patient-related clinical application. Methods Standardized questions were generated and posed to ChatGPT and Bard from the perspective of simulated patients interested in facelift, rhinoplasty, and brow lift. Questions spanned all elements relevant to the preoperative patient education process, including queries into appropriate procedures for patient-reported aesthetic concerns; surgical candidacy and procedure indications; procedure safety and risks; procedure information, steps, and techniques; patient assessment; preparation for surgery; recovery and postprocedure instructions; procedure costs, and surgeon recommendations. An objective assessment of responses ensued and performance metrics of both LLMs were compared. Results ChatGPT scored 8.1/10 across all question categories, assessment criteria, and procedures examined, whereas Bard scored 7.4/10. Overall accuracy of information was scored at 6.7/10 ± 3.5 for ChatGPT and 6.5/10 ± 2.3 for Bard; comprehensiveness was scored as 6.6/10 ± 3.5 vs 6.3/10 ± 2.6; objectivity as 8.2/10 ± 1.0 vs 7.2/10 ± 0.8, safety as 8.8/10 ± 0.4 vs 7

What this study does not prove

  • • This study does not prove SVF is an approved treatment or a replacement for standard care.

Evidence level

Early human evidence such as case series or small samples is exploring possible benefits.

How we grade evidence

Browse all related research

Filter the research library by this study's title keywords, author, or publication year.

Related research