scopusflow 0.3.0
A data and documentation release. The bundled corpus becomes a real harvest, and every example and article is rebuilt on it.
The bundled example records
The dataset the package ships for offline work was replaced outright.
-
example_recordsis now a worked example harvest of 138 real journal articles on graphene supercapacitors published between 2015 and 2024, carrying their real titles, DOIs, source titles, first authors and citation counts. It replaces the six invented records shipped previously. - The records are not a ‘Scopus’ harvest and are not described as one. The Elsevier API terms do not permit redistributing retrieved records, so they come from OpenAlex, whose metadata is released under CC0, reshaped into the schema [
scopus_fetch()] returns. The reasoning is recorded in the design notes. - The harvest is complete rather than sampled, so the rows per year are the real publications per year for that query. Eleven records carry no DOI and two no source title, kept as they arrive because a real harvest has such gaps.
-
scopus_idis empty throughout, these records not having come from ‘Scopus’, so de-duplication falls back to the DOI as it does for any record whose identifier is missing.
Documentation
The material a reader meets was rebuilt on the new corpus, and one misleading fixture was replaced.
- Every vignette and example runs on that corpus, paired with the key-gated live call a reader would actually write, and the figures quoted in the prose were recomputed against the new data.
- The demo mode of
run_app()draws on the same corpus, so a first look at the app shows real articles rather than invented rows. - The parser fixture in
inst/extdatamoves onto the reserved 10.5555 example prefix. It previously paired genuine, resolving DOIs with invented titles and authors, so a reader who checked one found a real paper mislabelled.
scopusflow 0.2.1
A documentation release. The vignettes now demonstrate several features that 0.2.0 shipped but did not show.
-
vignette("designing-queries")shows theAND NOToperator in [scopus_query()], excluding a dominant homonym from a search. -
vignette("analysing-a-literature")passes [scopus_intersections()] a concept that is already a complete field-tagged expression and so is used as given, letting a concept be a synonym set rather than a single term. -
vignette("scopusflow")introduces [scopus_top()] on a record set, and points to the analysis article for the plots and trends built on it. -
vignette("plans-and-quota")coversverbose = TRUEin [scopus_fetch_plan()], which reports a line as each cell is fetched or loaded from cache. -
vignette("building-a-reference-set")shows [scopus_extract_dois()] on a plain vector of DOIs, both with the default deduplication (which ignores case and resolver prefixes) and withdedupe = FALSE. -
vignette("keywords-and-references")tallies author keywords across a [scopus_corpus()] result, the per-keyword document count the article is named for. - [
scopus_abstract()]’s help page no longer describes the Python twin’s reference fields.
scopusflow 0.2.0
This release reaches further into the API, adds an analysis and export layer on top of a retrieval, and introduces a local, code-free app.
Deeper retrieval
A search now reaches further into the API, past the offset ceiling and beyond the fields the Search endpoint returns.
- [
scopus_fetch()] gainscursor = TRUE, cursor-based pagination that retrieves a whole large query without the 5000-record ceiling of offset paging. The warning on a query that exceeds the ceiling suggests this alongside partitioning with [scopus_plan()]. - [
scopus_fetch()] and [scopus_fetch_plan()] add anauthkeywordscolumn whenview = "COMPLETE"is requested, at no cost beyond that view’s own smaller page size. Theview = "STANDARD"output is unchanged, and [read_scopus_records()] keeps the column across a CSV round-trip. - [
scopus_abstract()] retrieves the abstract and fuller metadata for one or many records from the ‘Scopus’ Abstract Retrieval API, resilient to an identifier that cannot be found. Throughviewandinclude = c("references", "keywords")it also retrieves a document’s own reference list (as a structured, per-citation data frame, not a joined string) and author keywords, with per-identifier caching keyed by the requested view and extras, ann_requests/quotaattribute, and a clear, actionable error on an entitlement 403 that stops the batch rather than repeating the same failure for every identifier.include = "keywords"withoutview = "FULL"is rejected up front, since theREFresponse carries no author keywords. - [
scopus_corpus()] combines a search result with abstract retrieval into a minimalid/title/year/keywords/referencesshape for downstream tools such as keyword co-occurrence or citation-network analysis, without replacing [as_bibliometrix()]. A new vignette, Author keywords and references, walks through all of this with real DOIs. - [
scopus_fetch_plan()] compares each checkpoint’s recorded query with the plan cell before loading it, refetching and overwriting the checkpoint on a mismatch, so two different plans pointed at the samecache_dircannot serve each other’s records. A checkpoint that carries no query information (a zero-row cell, or one written by scopusflow 0.1.0) loads as before, and a cache directory is still best kept to a single plan.
Analysis and plots
The package gains a layer for summarising a literature and drawing the result.
- [
scopus_trend()] reports annual record counts for a query (the size of a literature over time), with [plot_scopus_trend()]. - [
scopus_top()] tallies the most frequent sources or authors in a record set, with [plot_scopus_top()], which draws whole-number axis breaks (so a tally of small counts shows no fractional ticks) and derives the count axis’s headroom from the widest end-of-bar label (so a wide count, say five figures on a top-authors bar, does not clip at the panel edge). Anautoplot()method draws a record set’s publications per year. - [
scopus_intersections()] counts a named set of concepts and any requested intersections of them, sizing where a study or a niche sits within the surrounding literature at one count request per row. Concept values that are already complete field-tagged expressions are used as given rather than wrapped again. [plot_scopus_intersections()] draws the result as a lollipop chart on a log-scale axis, with anautoplot()method and an optional highlight (for example the niche itself) whose legend label is derived from what is highlighted: ‘Focal intersection’ for intersections, ‘Focal concept’ for concepts and ‘Focal set’ for a mixture. An explicithighlight_labelstill wins. - [
plot_scopus_comparison()] gainslegend_inside. When set, and a legend is drawn, it is placed inside the panel in whichever corner has the most free space, on a small semi-transparent background, rather than above the panel. The default keeps the previous behaviour. - [
plot_scopus_comparison()] now spreads the direct line labels vertically so topics that converge near the final year no longer overlap, and falls back to a legend when there are too many topics to label legibly. The labels are spread when the figure is drawn, against the rendered text height, so they stay legible at any figure size, including a short panel such as the app’s result card. They carry no leader lines. The labels are colour-matched to their lines and spread in the same order as the line ends, so the link is clear without a leader that would otherwise cut across neighbouring labels.
Export
A retrieval can now leave the package in the formats other tools read.
- [
as_bibtex()] and [as_ris()] export a record set to the BibTeX and RIS interchange formats, so a search can be carried into Zotero, EndNote, Mendeley or a LaTeX bibliography.
A code-free app
The whole workflow is now available without writing any R.
- [
run_app()] launches a local, code-free Shiny app for building a search, retrieving records with a live progress terminal, and exporting them. A panel mirrors every choice as a runnable R script, so the app is an on-ramp to the package. It runs on your own machine, so the API key never leaves it. The app also has a Compare topics tab (with highlight, stability-band and counts-in-label controls, a per-term progress indicator, a quota estimate and a CSV export) and a Demo mode, on by default, that synthesises records and a comparison so the whole workflow can be explored with no key and no network. A new vignette, Using the code-free app, walks through every panel. - The app holds steady under stress. It refuses to start a comparison while a harvest is running, surfaces any comparison failure as a notification rather than a crash, floors a fractional maximum-records entry, drops duplicate comparison terms, and tells you when there is nothing to cancel.
scopusflow 0.1.0
CRAN release: 2026-06-20
First release.
- Reproducible search plans with [
scopus_plan()], and cheap sizing with [scopus_count()]. - Quota-aware, paginated retrieval through [
scopus_fetch()], with the largest page each view allows requested by default to keep request counts low, and resumable, cached, partitioned retrieval through [scopus_fetch_plan()]. - A stable normalised record schema from [
scopus_records()], with asummary()method that gives a quick overview. - DOI extraction and change tracking with [
scopus_extract_dois()] and [scopus_diff_dois()]. - Topic-trend comparison with [
scopus_compare_topics()], and a plot from [plot_scopus_comparison()] orautoplot(). - Interoperability and I/O through [
as_bibliometrix()], [write_scopus_records()] and [read_scopus_records()]. - A reference to the common ‘Scopus’ field tags in [
scopus_field_tags()], a safe query composer in [scopus_query()], and a bundled [example_records] dataset for offline exploration. - Safe merging of record sets with [
scopus_combine()] (and ac()method), plusas_tibble()andas.data.frame()coercion. - A typed condition system (
scopus_errorand its subclasses) and quota-header parsing with [scopus_quota()]. - The comparison plot uses whole-number year breaks, a colour-blind-safe palette, direct line labels, an optional
highlightargument and a shaded Wilson stability band (an illustrative range, switchable withinterval). - The bundled
example_recordsspans several disciplines, and the examples and five workflow vignettes draw on a wide range of fields. - Multiple authors are retained in the
authorscolumn rather than truncated to the first. Very large result totals are handled without overflow, and DOI cleaning copes withwww.doi.orghosts andDOI:labels.
