Repositorio Dspace

Accuracy of large language models in interpreting urological clinical guidelines: a comparative study with expert evaluation

Mostrar el registro sencillo del ítem

dc.contributor.author Borque-Fernando, Ángel
dc.contributor.author Navarro, Denis
dc.contributor.author Doblare, Manuel
dc.contributor.author Esteban, Luis-M
dc.contributor.author Pérez-Fentes, Daniel
dc.contributor.author Álvarez-Maestro, Mario
dc.contributor.author Medina-López, Rafael-A
dc.contributor.author Rodríguez-Faba, Óscar
dc.contributor.author Rubio-Briones, José
dc.contributor.author Fernández-Pello, Sergio
dc.contributor.author Fernández-Gómez, Jesus-María
dc.contributor.author Fernández-Aparicio, Tomás
dc.contributor.author Guerrero-Ramos, Felix
dc.contributor.author Izquierdo, Laura
dc.contributor.author Álvarez-Ossorio-Fernández, José-Luis
dc.date.accessioned 2026-08-03T10:26:54Z
dc.date.available 2026-08-03T10:26:54Z
dc.date.issued 2026-03
dc.identifier.issn 1756-2872
dc.identifier.uri https://sms.carm.es/ricsmur/handle/123456789/27157
dc.description.abstract BACKGROUND: Large language models (LLMs) are increasingly being explored to supporting evidence-based decision-making in urology, but their accuracy in interpreting and applying clinical guidelines remains uncertain. OBJECTIVES: We aimed to evaluate the ability of LLMs to interpret and apply clinical guidelines across the full spectrum of major urological cancers. DESIGN: This expert-validated study evaluated six configurations of three top LLMs (Claude, Gemini, and ChatGPT) using 25 structured questions for each of the seven major urological cancers: prostate cancer, upper tract urothelial carcinoma, muscle-invasive and non-muscle-invasive bladder cancer, renal cell carcinoma, penile cancer, and testicular cancer. METHODS: Both simple and rephrased prompts were used to assess the impact of prompt engineering on response quality. All figures and tables from the English-language EAU guidelines were systematically converted into plain, structured text and peer reviewed by multidisciplinary experts before evaluating the LLM responses. Each response was independently rated by 9-11 uro-oncology specialists using a five-point Likert scale (1: incorrect/unacceptable, 5: optimal), resulting in 10,500 evaluations. RESULTS: Claude achieved the highest overall accuracy, with 45.9% of responses rated as optimal (Likert 5) and 87% as optimal/acceptable (Likert 4-5). Tumor-specific performance peaked in muscle-invasive bladder (56.7% optimal, 93% optimal/acceptable), penile (49.5%, 95%), and testicular cancer (60.9%, 94%). Gemini and ChatGPT showed lower optimal rates but acceptable performance (68%-70% optimal/acceptable). Rephrased prompts did not consistently outperform simple versions. All models showed acceptable accuracy, but the results should be interpreted cautiously due to recency bias and fast LLM tech evolution. CONCLUSION: This study demonstrates the value of rigorous plain language adaptation and expert validation in benchmarking LLMs, supporting their potential as decision-support tools in uro-oncology.
dc.language.iso eng
dc.publisher SAGE PUBLICATIONS LTD
dc.rights Atribución/Reconocimiento-NoComercial 4.0 Internacional 
dc.rights.uri https://creativecommons.org/licenses/by-nc/4.0/deed.es  *
dc.title Accuracy of large language models in interpreting urological clinical guidelines: a comparative study with expert evaluation
dc.type info:eu-repo/semantics/article 
dc.identifier.pmid 41918915
dc.relation.publisherversion https://journals.sagepub.com/doi/10.1177/17562872261436905
dc.type.version info:eu-repo/semantics/publishedVersion 
dc.identifier.doi 10.1177/17562872261436905
dc.journal.title THERAPEUTIC ADVANCES IN UROLOGY
dc.identifier.essn 1756-2880


Ficheros en el ítem

Este ítem aparece en la(s) siguiente(s) colección(ones)

Mostrar el registro sencillo del ítem

Atribución/Reconocimiento-NoComercial 4.0 Internacional  Excepto si se señala otra cosa, la licencia del ítem se describe como Atribución/Reconocimiento-NoComercial 4.0 Internacional 

Buscar en DSpace


Búsqueda avanzada

Listar

Mi cuenta