|
View: |
Part 1: Document Description
|
|
Citation |
|
|---|---|
|
Title: |
Supplementary dataset and reproducible codes for LLM-assisted mapping feedstocks of eight conversion technologies from over 121,000 studies |
|
Identification Number: |
doi:10.18710/JM6U7B |
|
Distributor: |
DataverseNO |
|
Date of Distribution: |
2026-01-16 |
|
Version: |
2 |
|
Bibliographic Citation: |
Barahmand, Zahir, 2026, "Supplementary dataset and reproducible codes for LLM-assisted mapping feedstocks of eight conversion technologies from over 121,000 studies", https://doi.org/10.18710/JM6U7B, DataverseNO, V2 |
|
Citation |
|
|
Title: |
Supplementary dataset and reproducible codes for LLM-assisted mapping feedstocks of eight conversion technologies from over 121,000 studies |
|
Identification Number: |
doi:10.18710/JM6U7B |
|
Authoring Entity: |
Barahmand, Zahir (https://ror.org/05ecg5h20) |
|
Producer: |
University of South-Eastern Norway |
|
Date of Production: |
2025-11-26 |
|
Software used in Production: |
Python |
|
Software used in Production: |
Microsoft Excel |
|
Distributor: |
DataverseNO |
|
Distributor: |
University of South-Eastern Norway |
|
Access Authority: |
Barahmand, Zahir |
|
Depositor: |
University of South-Eastern Norway |
|
Date of Deposit: |
2025-12-16 |
|
Date of Distribution: |
2025-12-16 |
|
Holdings Information: |
https://doi.org/10.18710/JM6U7B |
|
Study Scope |
|
|
Keywords: |
Chemistry, Earth and Environmental Sciences, Engineering, bioeconomy, circular economy, biomass, conversion technologies, gasification, fermentation, pyrolysis, torrefaction, aerobic digestion, anaerobic digestion |
|
Abstract: |
This dataset was developed to systematically characterise feedstock–technology relationships across eight major biomass conversion technologies by mining a large Scopus-derived bibliographic corpus (1887–2025; partial coverage for 2025). The workflow is LLM-assisted and fully reproducible, combining automated extraction of feedstock and technology phrases from bibliographic text fields (titles, abstracts, and keywords) with rule-based cleaning and a subsequent LLM-based validation step, followed by targeted manual curation for final release. The dataset is intended for use in technology landscape analyses, evidence synthesis, and comparative assessments of biomass conversion pathways, where consistent and traceable feedstock descriptors are required across a very large volume of studies. A data descriptor titled "A large-scale, LLM-assisted and validated dataset of biomass and waste conversion technologies and feedstocks" with the following abstract will published based on this dataset: Biomass, organic wastes and biogenic by-products are increasingly targeted for low-carbon fuels and value-added chemicals. However, strategic decision-making from a circular economy perspective requires a big-picture view of the relative significance of different conversion technologies in handling diverse feedstock portfolios, and no large-scale, cross-technology mapping of these portfolios is currently available. Thus, a literature-derived dataset was assembled, that links eight major waste-to-x valorisation technologies (gasification, pyrolysis, hydrothermal liquefaction, torrefaction, anaerobic digestion, aerobic digestion, fermentation and transesterification) to their reported feedstocks. Using the Scopus database, 121,365 records were retrieved with harmonised search strings, spanning publications from 1887 to 2025. This constrained yet scalable search strategy both facilitates automated extraction and validation and yields a rich dataset. Further, a large language model assisted workflow was implemented to extract candidate technology and feedstock phrases, followed by a two-level validation that combines rule-based cleaning with targeted LLM re-evaluation to minimise manual curation. The resulting dataset provides technology-specific, validated feedstock descriptors that supports comparative analyses and decision-support applications in a circular bioeconomy context. |
|
Time Period: |
1887-2025 |
|
Date of Collection: |
2025-11-26-2025-11-26 |
|
Country: |
Norway |
|
Kind of Data: |
Data from Literature |
|
Methodology and Processing |
|
|
Sources Statement |
|
|
Data Sources: |
The data files in this dataset contain extracts from <a href="https://www.elsevier.com/en-gb/products/scopus/content" title="Scopus" target="_blank">Scopus</a>; reused under the <a href="https://www.elsevier.com/en-gb/legal/elsevier-website-terms-and-conditions" title="Elsevier" target="_blank">Elsevier Terms and Conditions</a>.<br><br> As of March 2025, Elsevier's Scopus database contains over 100 million records. The data files in this dataset contain extractions from approx. 121,364 Scopus records, which means approx. 0.12% of the Scopus database records. The extracted data thus only represent non-substantial portions of the Scopus database, and the dataset is not designed to replicate or replace the functionality of the Scopus database. Therefore, the reuse (including redistribution) of these extracts is permitted by the exceptions rules in IPR and database protection regulations, such as <a href=" http://data.europa.eu/eli/dir/1996/9/2019-06-06" title="Lawful users" target="_blank"> the EU Database Directive</a> (cf. article 8 Rights and obligations of lawful users), "uvesentlige deler av databaser" (Norway; cf. <a href="https://lovdata.no/lov/2018-06-15-40/§24" title="uvesentlige deler av databaser" target="_blank">§ 24 in Åndsverkloven</a>), "sitatretten" (Norway; cf. <a href="https://lovdata.no/lov/2018-06-15-40/§29" title="sitatretten" target="_blank">§ 29 in Åndsverkloven</a>). |
|
Data Access |
|
|
Notes: |
<a href="http://creativecommons.org/publicdomain/zero/1.0">CC0 1.0</a> |
|
Other Study Description Materials |
|
|
Label: |
0_ReadMe.txt |
|
Text: |
Detailed overview of the repository structure, file contents, workflow steps, reuse guidance, and notes on source traceability. The public release excludes Scopus-derived abstracts and raw Scopus export files. |
|
Notes: |
text/plain |
|
Label: |
Codes_Py1_Automated_data_extraction_llm.zip |
|
Text: |
Python scripts, configuration files, example inputs/outputs, and documentation for the LLM-assisted extraction of feedstock and technology descriptors from title and abstract fields. |
|
Notes: |
application/zip |
|
Label: |
Codes_Py2_Automated_extracted_data_cleaning.zip |
|
Text: |
Python scripts and documentation for rule-based cleaning, heuristic scoring, and validation-status assignment of the extracted feedstock and technology descriptors. |
|
Notes: |
application/zip |
|
Label: |
Codes_Py3_Automated_validation_llm.zip |
|
Text: |
Python scripts and documentation for targeted LLM-assisted validation of uncertain or non-accepted records after rule-based cleaning. |
|
Notes: |
application/zip |
|
Label: |
Review_DataExtraction_Validation_Curation_FinalDataset.zip |
|
Text: |
Processed and derived datasets from the extraction, cleaning, validation, and manual-curation workflow, including bibliographic source identifiers, intermediate outputs, and the final curated feedstock–technology descriptors. Scopus-derived abstracts and raw Scopus exports are not included. |
|
Notes: |
application/zip |