Repositorio Dspace

Annotation of biological samples data to standard ontologies with support from large language models

Mostrar el registro sencillo del ítem

dc.contributor.author Riquelme-García, Andrea
dc.contributor.author Mulero-Hernández, Juan
dc.contributor.author Fernández-Breis, Jesualdo-Tomás
dc.date.accessioned 2026-03-10T11:49:28Z
dc.date.available 2026-03-10T11:49:28Z
dc.date.issued 2025
dc.identifier.citation Riquelme-García A, Mulero-Hernández J, Fernández-Breis JT. Annotation of biological samples data to standard ontologies with support from large language models. Computational and Structural Biotechnology Journal. 2025;27:2155-67. doi:10.1016/j.csbj.2025.05.020
dc.identifier.issn 2001-0370
dc.identifier.uri https://sms.carm.es/ricsmur/handle/123456789/25233
dc.description.abstract The semantic integration of biological data is hindered by the vast heterogeneity of data sources and their limited semantic formalization. A crucial step in this process is mapping data elements to ontological concepts, which typically involves substantial manual effort. Large Language Models (LLMs) have demonstrated potential in automating complex language-related tasks and may offer a solution to streamline biological data annotation. This study investigates the utility of LLMs-specifically various base and fine-tuned GPT models-for the automatic assignment of ontological identifiers to biological sample labels. We evaluated model performance in annotating labels to four widely used ontologies: the Cell Line Ontology (CLO), Cell Ontology (CL), Uber-anatomy Ontology (UBERON), and BRENDA Tissue Ontology (BTO). Our dataset was compiled from publicly available, high-quality databases containing biologically relevant sequence information, which suffers from inconsistent annotation practices, complicating integrative analyses. Model outputs were compared against annotations generated by text2term, a state-of-the-art annotation tool. The fine-tuned GPT model outperformed both the base models and text2term in annotating cell lines and cell types, particularly for the CL and UBERON ontologies, achieving a precision of 47-64% and a recall of 88-97%. In contrast, base models exhibited significantly lower performance. These results suggest that fine-tuned LLMs can accelerate and improve the accuracy of biological data annotation. Nonetheless, our evaluation highlights persistent challenges, including variable precision across ontology categories and the continued need for expert curation to ensure annotation validity.
dc.language.iso eng
dc.publisher ELSEVIER
dc.rights Atribución/Reconocimiento 4.0 Internacional
dc.rights.uri https://creativecommons.org/licenses/by/4.0/deed.es
dc.title Annotation of biological samples data to standard ontologies with support from large language models
dc.type info:eu-repo/semantics/article
dc.identifier.pmid 40510764
dc.relation.publisherversion https://linkinghub.elsevier.com/retrieve/pii/S2001037025001837
dc.type.version info:eu-repo/semantics/publishedVersion
dc.identifier.doi 10.1016/j.csbj.2025.05.020
dc.journal.title Computational and Structural Biotechnology Journal


Ficheros en el ítem

Este ítem aparece en la(s) siguiente(s) colección(ones)

Mostrar el registro sencillo del ítem

Atribución/Reconocimiento 4.0 Internacional Excepto si se señala otra cosa, la licencia del ítem se describe como Atribución/Reconocimiento 4.0 Internacional

Buscar en DSpace


Búsqueda avanzada

Listar

Mi cuenta