R

Done with convergence warnings? Publications that used allFit to compare optimisers

Post

A convergence warning from lme4 does not by itself show that a model is wrong. One check that the lme4 authors recommend is to refit the model with every available optimiser through allFit and compare the estimates that matter. This post records the candidate publications found by a reproducible search of Scopus, Europe PMC and OpenAlex, and explains why the list can only be a starting point for screening.

How often do researchers say they speculate? Nine literatures compared

Post

This reproducible descriptive analysis uses annual Scopus counts to compare how often records in nine query-defined literatures contain a form of “speculate” in their titles, abstracts or keywords. The count is a proxy for explicit wording in indexed records and says little about the quality of the evidence.

depictr: One visual language from first look to final figure

Post

Figures often change their visual language as an analysis moves from exploration to reporting. This post uses depictr to examine a simulated lexical-decision experiment, plot its mixed-model estimates, audit a figure and correct the accessibility problems that the audit finds.

Greenhouse gas emissions in context: Raw totals, income groups and human development

Post

Greenhouse gas emissions can be reported as raw totals or per-capita values, each answering a different descriptive question. Using the 2025 release of the Emissions Database for Global Atmospheric Research, the World Bank income classification and the Human Development Index, this post compares totals, per-capita figures, sectors and a descriptive association with human development.

lexsync: From a word-frequency corpus to an EEG-ready experiment

Post

Selecting words is only the first step in preparing a psycholinguistic experiment. This post uses lexsync to match high- and low-frequency items, balance them across lists, and export timed PsychoPy and OpenSesame scripts that send electroencephalography (EEG) triggers, plus a jsPsych version that logs event markers.

pilotr: Pilot the study before running it

Post

A simulated semantic-priming study shows how precision, sign errors, magnitude errors and singular fits can be examined alongside power before choosing a design.

scopusflow: A literature search you can rerun

Post

A literature-search export does not preserve the decisions that produced it. This post uses scopusflow to define a search in advance, resume it from checkpoints and generate a PRISMA-S search record from the resulting data.

theoryforge: A theory you can check

Post

A theory written only in prose cannot be checked mechanically. This post represents a theory of modality switching as structured data, derives the conditional independencies it implies, and tests them against data simulated from the theory and from a rival account.

You shall know a word by the company it keeps — so choose your prompts wisely

Post

In 1957, the linguist J. R. Firth observed that ‘you shall know a word by the company it keeps’. The principle that words sharing contexts share meaning underlies generative AI, from early Latent Semantic Analysis to today's largest Transformers. This post traces that lineage with three interactive LSA-to-PCA visualisations in R, built from Reuters newswire, State of the Union addresses and IMDB reviews. They show where simple co-occurrence models succeed and where they fail, and the post explains how scale and the Transformer turned a modest insight into the technology behind ChatGPT. It then examines why LLMs, optimised for fluency, cannot be rid of hallucination entirely, and argues that careful prompt engineering is the best tool we have for steering a fundamentally heuristic machine.

R functions for checking and fixing vmrk files from BrainVision

Post

Electroencephalography (EEG) has become a cornerstone for understanding the intricate workings of the human brain in the field of neuroscience. However, EEG software and hardware come with their own set of constraints, particularly in the management of markers, also known as triggers. This article aims to shed light on these limitations and future prospects of marker management in EEG studies, while also introducing R functions that can help deal with vmrk files from BrainVision.

rscopus_plus: An extension of the rscopus package

Post

Extension of the rscopus R package with functions that manage search quotas, retrieve DOIs for reference managers, search for additional DOIs, compare publication counts across topics, and visualize bibliometric comparisons over time.

FAQs on mixed-effects models

Post

Frequently asked questions about mixed-effects models, covering the necessity of random slopes, appropriate p-value calculation methods, parallelization limitations, convergence issues, and optimizer selection.

FAIR standards for the creation of research materials, with examples

Post

In the fast-paced world of scientific research, establishing minimum standards for the creation of research materials is essential. Whether it's stimuli, custom software for data collection, or scripts for statistical analysis, the quality and transparency of these materials significantly impact the reproducibility and credibility of research. This blog post explores the importance of adhering to FAIR (Findable, Accessible, Interoperable, Reusable) principles, and offers practical examples for researchers, with a focus on the cognitive sciences.

Preprocessing the Norwegian Web as Corpus (NoWaC) in R

Post

An R script for preprocessing frequency list data from the Norwegian Web as Corpus (NoWaC), including instructions for downloading and preparing the corpus data.

ggplotting power curves from the simr package

Post

A custom R function to create ggplot2 visualizations of power curves generated by the simr package's powerCurve function for mixed-effects models.

How to discretise the colour variable in sjPlot::plot_model into equally-sized intervals

Post

Custom functions that extend sjPlot to discretise continuous variables in interaction plots into equally-sized intervals (deciles or sextiles) that include minimum and maximum values, with legend counts showing sample sizes per level.

How to map more informative values onto fill argument of sjPlot::plot_model

Post

Custom function that extends sjPlot to display original categorical variable values in interaction plot legends instead of transformed values, facilitating clearer communication when using sum-coded or other transformed predictors.

How to visually assess the convergence of a mixed-effects model by plotting various optimizers

Post

A custom R function to create ggplot2 visualizations of fixed effects from models refitted with multiple optimizers using lme4's allFit function, enabling visual assessment of convergence validity in mixed-effects models.

Table joins with conditional “fuzzy” string matching in R

Post

Implementation of fuzzy string matching in R using fuzzyjoin::stringdist_join with controlled fuzziness via max_dist argument, demonstrated with food name matching and frequency-based code selection.

A new function to plot convergence diagnostics from lme4::allFit()

Post

When a model has struggled to find enough information in the data to account for every predictor—especially for every random effect—, convergence warnings appear (Brauer & Curtin, 2018; Singmann & Kellen, 2019). In this article, I review the issue of convergence before presenting a new plotting function in R that facilitates the visualisation of the fixed effects fitted by different optimization algorithms (also dubbed optimizers).