How well is the field defining its constructs, using theory generatively rather than rhetorically, and avoiding the conflation of distinct concepts?
Week 2 argues that a claim must be defined precisely enough to be tested, that abstract constructs must be distinguished from the measures that index them, and that a theory earns its place by generating specific, falsifiable predictions rather than by lending rhetorical legitimacy. The empirical questions are whether the field uses theory generatively, defines its constructs consistently, and clarifies terms when they become ambiguous.
The most direct evidence comes from a large-scale computational audit conducted by the SSRC (Warne, Gorman et al., in preparation; unpublished). Using an automated, deterministic large-language-model pipeline to analyse the introductions and discussions of 269 applied sports-science articles, the audit identified 421 theoretical-framework instances (mean 1.6 per article; 93% explicitly named). Only 9% of these theories were classified as operational — directly generating the specific prediction the study tested. Strong hypothesis–theory alignment, in which all hypotheses were derived from the stated theory rather than merely stated as consistent with it, occurred in just 1% of articles, and full re-engagement with the theoretical framework in the discussion in only 5%. The authors conclude that theory use in the field is “more rhetorical than generative.” (As this is an in-preparation working paper, figures should be treated as provisional pending peer review.)
The SSRC's FAIR Theory Database — an internal, machine-readable catalogue of named theories in sport and exercise science — scores each theory on operationalisation clarity, replication support, cumulative development, precision/falsifiability, and domain breadth (a Lakatosian and Meehlian framework). Across its 97 catalogued theories, 36 reach the highest tier while 48 sit in the middle tier and 13 in the lowest — that is, roughly 63% show gaps in operationalisation, replication, or falsifiability that place them below the top tier. This is an internal resource rather than a peer-reviewed dataset, but it is consistent with the audit above: many named theories in the field are not yet developed to the point of generating precise, testable predictions.
The broader open-science literature reaches the same conclusion from a different angle. Van Lissa and colleagues (2026) argue that a decade of reform has improved the “back end” of research (data, code, analysis) while leaving theory specification — the front end — informal, verbal and non-interoperable, so that the same theoretical label routinely covers different operationalisations across studies. In health-behaviour research, Prestwich and colleagues (2014) found that only 56% of 190 physical-activity and dietary interventions reported any theory base, and that where theory was named it was often applied loosely rather than used to select intervention components.
Several prominent constructs lack agreed definitions or measures. After roughly three decades of research, mental toughness still shows persistent disagreement over whether it is uni- or multi-dimensional, a trait or a state, and how it should be operationalised (Gucciardi, Hanton, Gordon, Mallett & Temby, 2015). A clear jangle fallacy comes from the adjacent achievement literature: Credé, Tynan and Harms (2017), meta-analysing 88 samples (66,807 individuals), found “grit” correlates with conscientiousness at r = 0.84 — strong enough to question whether it is distinct at all.
The field also contains clear positive examples. “Fatigue,” long used as one word for distinct phenomena, was decomposed by Enoka and Duchateau (2016) into performance and perceived fatigability. “Training load” has been repeatedly clarified by Impellizzeri, Marcora and Coutts (2019) — external versus internal load — and extended by Impellizzeri and colleagues (2023) using exposure–dose concepts from epidemiology. Distinguishing constructs from their measures, however, is still frequently neglected: Flake and Fried (2020) document widespread questionable measurement practices, and Impellizzeri and Marcora (2009) note that surrogate and field measures are often adopted without validation (developed in Week 5).
| Source | Finding | Implication |
|---|---|---|
| Warne, Gorman et al. (in prep.) SSRC working paper — unpublished | Audit of 269 articles: 9% of theories operational; 1% strong hypothesis–theory alignment; 5% discussion re-engagement | Theory use is rhetorical rather than generative |
| FAIR Theory Database SSRC internal catalogue | 97 theories scored; ~63% below the top development tier (48 mid, 13 low) | Theoretical development in the field is uneven |
| Van Lissa et al. (2026) Persp. Psych. Science (adjacent) | Theory specification remains informal and non-interoperable despite open-science reform | Same label covers divergent operationalisations |
| Prestwich et al. (2014) Health Psychology | Only 56% of 190 PA/diet interventions reported a theory base; theory applied loosely | Theory use is inconsistent |
| Gucciardi et al. (2015) J. Personality | No consensus on dimensionality, traitness or measurement of mental toughness | A popular construct remains poorly defined |
| Credé, Tynan & Harms (2017) JPSP (adjacent) | “Grit” correlates with conscientiousness at r = 0.84 (88 samples) | Jangle fallacy: a “new” construct duplicates an old one |
| Enoka & Duchateau (2016) MSSE | Decomposed “fatigue” into performance vs perceived fatigability | Positive example: resolving a jingle fallacy |
| Impellizzeri, Marcora & Coutts (2019); Impellizzeri et al. (2023) IJSPP; Sports Medicine | Re-clarified and refined external vs internal training load | Positive example: active conceptual clarification |
| Flake & Fried (2020); Impellizzeri & Marcora (2009) AMPPS; IJSPP | Measures used without validity evidence; scores treated as constructs | Construct–measure conflation is common |
The direct audit is the decisive evidence: if only about 9% of cited theories generate a specific tested prediction and only 1% of articles show full hypothesis–theory alignment, then the field is largely not using theory in the generative sense the week describes. This is reinforced by an internal catalogue showing uneven theoretical development, by adjacent evidence that theory specification remains informal, and by persistent definitional problems for popular constructs. Because conceptual and theoretical clarity is upstream of measurement, design and analysis, these weaknesses propagate into the replication problems documented in Week 1.
Two qualifications keep the assessment proportionate. First, the strongest single figure comes from an unpublished, in-preparation audit and should be treated as provisional until peer-reviewed, though it is consistent with independent lines of evidence. Second, the field contains clear, successful examples of conceptual clarification (fatigue, training load), showing the problem is tractable and requires rigorous definition rather than additional resources.
Credé, M., Tynan, M. C., & Harms, P. D. (2017). Much ado about grit: a meta-analytic synthesis of the grit literature. Journal of Personality and Social Psychology, 113(3), 492–511. https://doi.org/10.1037/pspp0000102
Enoka, R. M., & Duchateau, J. (2016). Translating fatigue to human performance. Medicine & Science in Sports & Exercise, 48(11), 2228–2238. https://doi.org/10.1249/MSS.0000000000000929
FAIR Theory Database (2026). Internal machine-readable catalogue of theories in sport and exercise science (v1.3, 97 theories). Sports Science Replication Centre / OpenClaw Research Ecosystem. Internal resource.
Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456–465. https://doi.org/10.1177/2515245920952393
Gucciardi, D. F., Hanton, S., Gordon, S., Mallett, C. J., & Temby, P. (2015). The concept of mental toughness: tests of dimensionality, nomological network, and traitness. Journal of Personality, 83(1), 26–44. https://doi.org/10.1111/jopy.12079
Impellizzeri, F. M., & Marcora, S. M. (2009). Test validation in sport physiology: lessons learned from clinimetrics. International Journal of Sports Physiology and Performance, 4(2), 269–277. https://doi.org/10.1123/ijspp.4.2.269
Impellizzeri, F. M., Marcora, S. M., & Coutts, A. J. (2019). Internal and external training load: 15 years on. International Journal of Sports Physiology and Performance, 14(2), 270–273. https://doi.org/10.1123/ijspp.2018-0935
Impellizzeri, F. M., Shrier, I., McLaren, S. J., et al. (2023). Understanding training load as exposure and dose. Sports Medicine, 53(9), 1667–1679. https://doi.org/10.1007/s40279-023-01833-0
Prestwich, A., Sniehotta, F. F., Whittington, C., Dombrowski, S. U., Rogers, L., & Michie, S. (2014). Does theory influence the effectiveness of health behavior interventions? Meta-analysis. Health Psychology, 33(5), 465–474. https://doi.org/10.1037/a0032853
Van Lissa, C. J., Peikert, A., Ernst, M. S., van Dongen, N. N. N., Schönbrodt, F. D., & Brandmaier, A. M. (2026). To be FAIR: theory specification needs an update. Perspectives on Psychological Science. https://doi.org/10.1177/17456916251401850
Warne, J., Gorman, B., et al. (in preparation). Theory use in applied sports science: a computational audit of explicit and implicit theoretical frameworks, hypothesis derivation, and discussion re-engagement. Sports Science Replication Centre working paper, Technological University Dublin. Unpublished.