The pipeline extracted nearly 9,000 records covering more than 100 concrete-material attributes in about an hour.
The pipeline uses large language models to identify and structure information about concrete composition, processing and properties in otherwise unstructured scientific literature. The researchers report that it performed well across a broad range of language models, with an F1 score of up to 0.98 for different data attributes.
In about an hour, the system extracted nearly 9,000 high-quality records containing more than 100 attributes. Machine-learning analyses using the database found that larger, more diverse and more information-rich datasets improved predictions for both familiar materials and materials not represented in the training data.
What the AI extracted
The researchers developed a general-purpose pipeline for extracting materials data from scientific papers, using concrete as a demanding test case. It achieved an F1 score of up to 0.98 across composition, processing and property attributes and worked robustly across a broad range of large language models.
From a corpus screened from more than 27,000 publications, the pipeline produced nearly 9,000 high-quality records with more than 100 attributes in about one hour. These records were used to build what the researchers describe as the largest open laboratory database for blended cement concrete. Their machine-learning analyses also showed that larger, more varied and more information-rich datasets improved both accuracy on familiar cases and generalization to unseen materials.
Why the database matters
Experimental materials data are often scattered through papers and difficult to reuse. By organizing information from the literature into a searchable open database, the pipeline could make more concrete data available for data-driven research and reduce the manual work involved in building such datasets.
The findings also point to the value of dataset size and diversity when machine-learning models are used to study concrete materials. The researchers say the pipeline can be adapted to other materials fields, although this study demonstrates it directly for concrete.
Evidence and caveats
This is a journal article reporting the development and testing of an automated data-extraction pipeline, along with machine-learning analyses of the resulting database. The reported F1 score, a measure that combines how accurately the system finds relevant information and how completely it captures it, reached 0.98 for some attributes. The abstract does not specify the evaluation data, the performance of each attribute, or the types and frequency of extraction errors.
The database focuses on blended cement concrete, and the abstract does not establish how well the pipeline performs in other materials domains. It also does not describe the publication-screening criteria, the full coverage of the literature or whether all extracted records are equally complete. The claim that the pipeline is adaptable beyond concrete is therefore a proposed capability rather than a result demonstrated in this study.
// Source
npj Computational Materials · 2026 · DOI: 10.1038/s41524-026-02304-6
Authors: Zhanzhao Li, Kengran Yang, Qiyao He, Kai Gong
Institutions: Princeton University, Rice University, Kennedy Center