Everything on this page is a preprint: a manuscript that has been submitted to a journal but has not yet completed peer review. These are author versions. They may contain errors, their findings may change substantially during review, and they should not be treated as established evidence or used to guide clinical or policy decisions.
Where a manuscript is later published, the peer-reviewed version of record supersedes the file here and is linked from the Publications page. Please cite the version of record whenever one exists.
I post preprints because I think work should be readable while it is being judged, not only after. It is the same reasoning behind FAIR Press Journals, which I co-founded: research is more useful when the reasoning, the code and the data are open to inspection.
Normality Testing in Physiotherapy Research: A Full-Text Bibliometric Audit of a Common but Low-Stakes Error
Normality tests such as Shapiro–Wilk or Kolmogorov–Smirnov are widely used to choose between parametric and non-parametric analyses — but the assumption they concern applies to the residuals of the fitted model, not to the raw outcome data. This full-text bibliometric audit covers articles published in 16 physiotherapy and rehabilitation journals between 2015 and 2024 that reported a named normality test. Of 1,155 definitively codable papers, correcting for classifier labelling error gives a conservative prevalence of 89.0 % (95 % CI 81.4–95.2) applying the test to raw data rather than residuals, rising to 90.4 % among determinable papers; only 12 papers (1.0 %) tested residuals. The error exceeded 85 % in every journal with at least ten coded papers and showed no material change over the decade. A simulation across four distributions and sample sizes from 10 to 100 per group shows the power cost never exceeded about three percentage points — the test eventually chosen mattered far more than the residual-versus-raw distinction. Misapplication is the norm, the consequences are small, and correct practice is essentially free.
Does the choice of Sleep Regularity Index software change the conclusions? A re-analysis of the NHANES 2011–2014 sleep regularity–obesity association
The Sleep Regularity Index can be computed by several open-source packages, and the package
alone is known to change SRI values and their apparent associations with health outcomes.
I test whether a published SRI–obesity finding survives the choice of calculator, and why.
On the byte-identical minute-level series of my own published sample (n = 7,085) I recomputed
the SRI with GGIR and sleepreg at defaults and compared them with the
published hand-computed values. The calculators disagreed: package SRIs were about 7 median
points lower, correlated 0.81 and 0.67 with the published SRI, and reshuffled about half of
participants across quintiles. Every calculator preserved the direction of the effect, but not
its magnitude — the fully adjusted top-versus-bottom-quintile effect on BMI shrank from about
−5 % to −1.4 %. The divergence was structured and confined to incompletely recorded
participants: among the 60 % with seven consecutive days the effect was calculator-invariant,
whereas the entire calculator dependence arose in the 40 % with a day gap — an instance of
regression dilution from exposure measurement error. The qualitative finding is robust to the
calculator; its effect size and the sex/ethnic effect modification are not.
Moving the goalposts on purpose: almost sure hypothesis testing and what it would do to a physiotherapy trial
Almost every clinical trial still declares its result at a single fixed threshold, a p-value below 0.05. The line does not move whether the sample is small and fragile or enormous, and a classical result shows such a fixed line can never be consistent: under a true null the standardised statistic crosses 1.96 infinitely often, however much data is collected. I examine almost sure hypothesis testing, in which the significance level shrinks with the sample size (αn = n−p, p > 1), and ask what it does to the interpretation of a real physiotherapy trial. Simulation under a true null establishes that a fixed-α rule accumulates false rejections without bound while the shrinking level does not. I then tabulate the rejection threshold the method implies across the sample sizes physiotherapy trials actually reach, and apply it in a fully reproducible secondary re-analysis of the reported summary statistics of a registered, open-access randomised trial.
Meta-science in physiotherapy research: a structured review of what we know and what we still do not
Physiotherapy has built a large clinical evidence base, but whether that evidence can be trusted depends not only on the trials inside it — it depends on how the research is conducted, reported, and shared. Meta-research makes those questions answerable, yet its findings for physiotherapy had never been drawn together. This structured review maps physiotherapy-specific meta-research across the recognised themes of the field (methods, reporting, reproducibility, evaluation, incentives), charts the evidence by domain and rates each domain's maturity. The picture is uneven: trial quality has risen slowly; nine in ten musculoskeletal reviews are critically low on AMSTAR 2; spin appears in close to nine in ten "non-significant" trial reports and shifts clinicians' interpretation; baseline characteristics are tested for "significance" in three-fifths of recent trials; only one-third of trials are usefully registered; and replication is essentially absent while open data are rare. Conflicts of interest, citation bias, peer-review quality, the incentive and evaluation systems, and the use of large language models remain largely or wholly unstudied.