scopusflow: A reproducible workflow layer for Scopus bibliographic searches

Software

Abstract

Twin R and Python packages for planning, running and documenting reproducible Scopus searches. scopusflow handles quotas and retries, supports resumable retrieval, normalises records, tracks changes in DOI sets and produces PRISMA-S search records.

Overview

scopusflow treats a bibliographic search as a versioned workflow rather than a one-off export. A plan records the query and partitions, completed cells are cached for safe resumption, and the resulting object retains retrieval metadata, package version and DOI-level changes. The R and Python packages use the same workflow concepts while respecting the access and redistribution limits of the source database.

Illustrative output

The example below uses the package's offline stand-in corpus, so the plot demonstrates the workflow rather than claiming to describe the live Scopus literature.

Line chart showing the number of graphene-supercapacitor records in an offline demonstration corpus from 2015 to 2024
Records per year in the reproducible offline demonstration.

Reproducible reporting

Keep the plan, cache manifest, native result object, DOI comparison and generated PRISMA-S record under version control. A future update can then distinguish a change in the literature from a change in the query, database coverage or retrieval process. Read the R documentation, Python documentation and companion blog post for the offline and authenticated workflows.

Reference

Bernabeu, P. (2026). scopusflow: A reproducible workflow layer for Scopus bibliographic searches (Version 0.4.0) [Computer software]. CRAN. https://doi.org/10.32614/CRAN.package.scopusflow

Comments are provided by Disqus and are not loaded automatically. Loading them connects your browser to Disqus, which may use cookies and process data under its privacy policy. See this site's privacy notice for details.