Purpose: This study aimed to compare the quality of audiological case interpretation and management recommendations generated by two freely-accessible large language model (LLM) chatbots (ChatGPT 3.5 and Gemini 2.0 Flash) with those of qualified audiologists, using an expert-rated quality assessment tool.
Method: This study followed a two-phased, cross-sectional design with comparative analysis: Phase 1: 15 anonymised adult audiological cases were interpreted by ten qualified audiologists (>2 years’ experience) and two LLM chatbots, producing 60 interpretations; each interpretation answered two questions (diagnostic interpretation and management), giving 120 responses. . Phase 2: 30 qualified audiologists independently rated all 120 responses using the CLEAR tool (5 raters per response; 600 total ratings), blinded to responder identity.
Results: There were no statistically significant differences in CLEAR scores between AI-generated and audiologist-generated responses (all p > 0.24 in mixed-effects modeling). Absolute differences did not exceed 1.2 CLEAR points on a 5-25 scale, with 95% confidence intervals crossing zero, indicating no statistically detectable differences in expert-rated quality. Inter-rater reliability was poor for single audiologist raters (ICC = 0.17-0.18) but moderate for aggregated scores (ICC = 0.50-0.53). Internal consistency was excellent (Cronbach’s alpha = 0.94-0.95).
Conclusions: Under controlled conditions, no statistically significant differences in CLEAR scores were detected between LLM-generated and audiologist-generated responses. These findings do not establish formal equivalence, but represent a meaningful empirical benchmark for this emerging area of research.