The pipeline uses large language models to identify and structure information about concrete composition, processing and properties in otherwise unstructured scientific literature. The researchers report that it performed well across a broad range of language models, with an F1 score of up to 0.98 for different data attributes.

In about an hour, the system extracted nearly 9,000 high-quality records containing more than 100 attributes. Machine-learning analyses using the database found that larger, more diverse and more information-rich datasets improved predictions for both familiar materials and materials not represented in the training data.