
LifeDive is a cross-kingdom phenotype resource containing more than seven million standardized estimates for 35,856 animal, plant, and fungal species, generated with a web-search-enabled large language model. Validation against independent databases and a detailed audit suggests that the dataset preserves strong biological structure and can support exploratory research, comparative analyses, and hypothesis generation.
The manuscript describing the data and methods is here, and all code for using the pipeline for your own work is here. Tyler can be contacted at tymoore@pennmedicine.upenn.edu.
Download the full data set here.
Suggested citation:
Moore, T. M. (2026). A language model generates seven million phenotype estimates across the tree of life. Retrieved from https://lifedive.org.
An interface for exploring the database is available below.