Psychology has spent much of the past decade discussing its replication problem as though it were primarily a problem of method.
Small samples. Low statistical power. Researcher degrees of freedom. Selective reporting. HARKing. Publication bias. The almost ritual importance once attached to p < .05.
All of these matter.
The Open Science Collaboration's attempt to replicate 100 published psychology studies remains difficult to ignore. Ninety-seven per cent of the original studies reported statistically significant results. Thirty-six per cent of the replications did. Replication effect sizes were, on average, around half the size of the originals (Open Science Collaboration, 2015).
That does not mean that 64 per cent of psychology was "wrong". The original paper itself used several measures of replication success, which produced different answers. Reducing the project to a league table of successes and failures rather misses the point.
Something was clearly wrong, though, and I am not convinced that another adjustment to the research methods handbook gets us very far.
I increasingly wonder whether we have concentrated too heavily on correcting researchers while paying rather less attention to the environment in which they were working.
What were we rewarding?
Imagine two early-career psychologists.
One proposes a careful, adequately powered replication of an important finding. It might produce a null result.
The other has a novel hypothesis, several potentially interesting comparisons and a reasonable chance of producing a statistically significant finding around which a decent story can be constructed.
For much of modern academic psychology, which was the better career bet?
Publications help careers. Grants help careers. Citations help careers. Research income is visible. Papers can be counted. Journal rankings can be recorded. Successful funding applications sit nicely on CVs and institutional dashboards.
The same logic reaches beyond research. Courses created, student numbers, projects delivered and external partnerships can all be counted.
Careful work establishing that an existing psychological effect is smaller than previously thought, dependent upon particular conditions, or perhaps absent altogether is rather harder to turn into an impressive number on a spreadsheet.
I don't think this requires us to imagine armies of psychologists consciously manipulating science to further their careers. That explanation is too easy and, I suspect, usually wrong.
Smaldino and McElreath (2016) offered a much more interesting possibility in The Natural Selection of Bad Science. Their modelling showed how research practices that generate more publishable findings can spread when publication success contributes to career success. No deliberate cheating is required.
That ought to make psychologists slightly uncomfortable.
We spend considerable time teaching students that behaviour is shaped by reinforcement, contingencies and environmental pressures. Yet academia has sometimes talked about scientists as though they somehow stand outside those principles.
Nosek and colleagues (2015) made a related point when developing the Transparency and Openness Promotion guidelines. Scientists may value openness and reproducibility while working within reward structures that do not adequately encourage either. An emphasis on innovation can work against verification.
Perhaps replication failure was partly telling us what our academic environment had been selecting for.
What were journals rewarding?
There is another participant in this story.
The journal.
A researcher cannot create publication bias alone. Somebody has to decide what gets published.
Editors and reviewers have exercised enormous influence over what eventually became the psychological literature. Novel findings were interesting. Significant findings were interesting. Clean stories were attractive. Replications were often less so. Null findings had an unfortunate tendency to disappear.
Munafò and colleagues (2017), in their manifesto for reproducible science, explicitly placed publication bias and incentive structures within the reproducibility problem. Positive, novel and tidy findings have historically been easier to publish than negative findings, replications and studies containing awkward loose ends.
So there is a question worth asking.
When an elegant, novel, statistically significant but ultimately fragile finding appeared in a prestigious journal, who exactly had succeeded?
The author had.
But the reviewers had recommended it. The editor had selected it. The journal had published it.
That doesn't make any of them villains. Editors and reviewers were operating within incentives of their own.
That is the point.
The replication crisis cannot simply be a story about researchers using inadequate statistics. It is also a story about what the machinery of academic publishing learned to recognise as good science.
What should we teach differently?
There is a danger that open science simply becomes another collection of procedures added to the research methods curriculum.
Students learn about statistical power. They preregister something. Perhaps they upload materials to the Open Science Framework. We explain p-hacking and HARKing and tell them not to do either.
Useful, certainly.
But I think the lesson from replication failure is more fundamental.
Students need to learn that uncertainty is not an embarrassment to be edited out of the final report.
A null result tells us something. A failed prediction tells us something. An effect which becomes substantially smaller when the sample becomes larger tells us something.
And a carefully conducted replication producing a different result has not necessarily "failed".
Perhaps research methods should therefore be taught less as a machinery for obtaining an answer and more as a way of giving our ideas a serious opportunity to be wrong.
Students should encounter messy data. They should see how reasonable analytical choices can alter conclusions. They should attempt replications and interrogate published papers rather than treating publication as a certificate of truth.
Occasionally, they should be rewarded for discovering absolutely nothing exciting.
A carefully designed project providing good evidence that something probably does not happen may demonstrate considerably more scientific maturity than an underpowered project containing three significant p-values.
Registered Reports tell us something interesting
Registered Reports are one of the more convincing responses to the problem.
The research question and methodology are reviewed before the results are known. In-principle acceptance therefore makes eventual publication much less dependent upon whether the results are significant or surprising.
Soderberg and colleagues (2021) compared 29 published Registered Reports with 57 conventional comparison papers. The Registered Reports received substantially higher assessments for methodological and analytical rigour, while ratings of novelty and creativity were statistically indistinguishable.
That is encouraging.
But I find the existence of Registered Reports interesting for another reason.
Why did we need to invent a special publication format to make sure that a well-designed study remained publishable when its hypothesis wasn't supported?
Registered Reports help solve a problem, but they also tell us something about the publishing system that made them necessary.
Universities have the same problem
The argument doesn't stop at journals.
Universities have become remarkably good at counting things: publications, citations, research income, grant applications, impact, student numbers, courses developed, projects delivered, external partnerships.
The attraction is obvious. These things can be recorded, compared and aggregated.
But what is measurable and what is valuable are not necessarily the same thing.
The research community increasingly recognises this. The Coalition for Advancing Research Assessment (CoARA) argues that research assessment should rely primarily on qualitative judgement, with quantitative indicators used responsibly. It explicitly challenges inappropriate reliance on Journal Impact Factor, h-index and simple counts of things such as papers, citations and grants as proxies for quality.
It also argues for recognition of activities that metric-heavy systems can overlook: teaching, supervision, mentoring, teamwork, technical roles and contributions to openness and research robustness.
There is an obvious parallel with education.
Some of the most valuable things educators do are remarkably difficult to count.
Noticing that somebody hasn't understood something. Changing an explanation because a room has gone quiet. Spending twenty minutes unpicking an argument. Recognising that a student is struggling before they tell you. Helping somebody become intellectually confident enough to disagree.
These things do not travel particularly well through dashboards.
The danger comes when the measurable activity gradually becomes the valued activity.
Research is susceptible to exactly the same process.
Then along came AI
Now we have another apparent crisis in academic publishing: "AI slop".
There are good reasons to be concerned.
Fabricated references are unacceptable. Invented evidence is unacceptable. Submitting machine-generated material that the named author neither understands nor checks is unacceptable. A reviewer who hands intellectual judgement over to a language model and returns its output under their own name has misunderstood the job.
Confidentiality creates another problem. Unpublished manuscripts cannot simply be fed into public AI systems without considering what happens to that material.
I don't find any of that particularly controversial.
What interests me more is the suggestion, sometimes lurking beneath the current discussion, that AI has arrived and contaminated an otherwise healthy scholarly ecosystem.
That version of history is harder to accept.
Peer review was inconsistent before ChatGPT.
Publication bias existed.
Prestige hierarchies existed.
Reviewers misunderstood papers.
Authors wrote formulaic prose.
Editors selected novelty.
Academics were under pressure to publish.
AI didn't create those conditions.
It entered them.
AI as a stress test
This is increasingly how I think about generative AI and academic publishing.
AI is stress-testing the system.
A stress test doesn't necessarily create the weakness. It can expose one that was already there.
Large language models can produce an academic-looking introduction, a conventional limitations section, plausible reviewer comments and fluent scholarly prose very cheaply.
That is clearly a problem when volume replaces thought.
But it raises another question.
If a machine can reproduce many of the surface characteristics by which we have learned to recognise scholarship, how reliable were those characteristics as indicators of scholarship in the first place?
Perhaps the unsettling thing about AI-generated academic prose isn't simply that machines have become good at imitating academics.
Perhaps some academic writing had become unusually easy to imitate.
The same applies to reviewing.
A generic review containing familiar comments about sample size, generalisability, theoretical contribution and the need to expand the discussion can sound impressively academic while requiring remarkably little engagement with the paper.
Humans were capable of producing reviews like that long before ChatGPT.
AI simply makes them cheaper, faster and scalable.
The evidence is awkward
Recent evidence makes a straightforward "human good, AI bad" account difficult to sustain.
A 2026 analysis by the Organization Science AI Task Force reported 42 per cent more submissions when comparing its specified pre- and post-ChatGPT periods. Writing quality also declined, and the authors' analyses attributed much of these changes to heavily AI-generated submissions. AI-generated writing in reviews was also associated with lower writing quality and less topical diversity.
That is serious evidence of a problem.
But another 2026 study makes the picture more complicated.
Thakkar and colleagues conducted a randomised study involving more than 20,000 reviews at ICLR 2025. Reviewers received AI-generated feedback about weaknesses in their own reviews. Twenty-seven per cent subsequently revised them, incorporating more than 12,000 suggestions. Blinded evaluation found the revised reviews more informative.
The AI wasn't reviewing the paper instead of the reviewer.
It was helping the reviewer review.
That distinction matters.
AI therefore appears capable of contributing to poor reviewing. Under different conditions it can also improve reviewing.
Which suggests that "Was AI involved?" is not a particularly useful final question.
A better one might be:
Where was the judgement?
A human can produce slop. I have, and continue to do so. Perhaps you think this piece is. Ha!
AI can produce slop extraordinarily efficiently.
But a knowledgeable person can also use AI to challenge an argument, identify something they have missed, test an alternative interpretation or improve the clarity of a review while remaining responsible for every judgement eventually made.
Perhaps AI is exposing something
I wonder whether a small part of the hostility towards AI in academic publishing is defensive.
Only a small part. The problems of confidentiality, fabricated material, provenance, accountability and industrial-scale low-quality submissions are real.
But AI exposes something else.
It shows how much scholarly production had already become formulaic.
The replication crisis challenged the assumption that publication was necessarily a reliable proxy for truth.
Generative AI is now challenging another assumption: that academic-looking output is necessarily a reliable proxy for intellectual contribution.
Both problems concern proxies.
There is a risk that "AI slop" becomes a convenient boundary marker. Humans inside. Machines outside. Scholarship on one side, synthetic contamination on the other.
That would be reassuring.
It would also save the existing publishing system from examining how much poor, formulaic or superficial material humans were already perfectly capable of producing.
There is a wonderfully awkward example here.
Giray and colleagues published a 2026 paper specifically examining "AI slop in academic publishing". A corrigendum was subsequently required because six references contained incorrect DOI information. The cited publications themselves were genuine and the errors did not alter the paper's arguments.
I don't mention this as a cheap gotcha against the authors. Quite the opposite.
It is a useful reminder of why the boundary between trustworthy human scholarship and unreliable machine output is nowhere near as clean as we might like it to be.
Humans make mistakes.
Machines make mistakes.
The answer in either case is scrutiny.
Two stress tests
This is where replication and generative AI unexpectedly meet.
Replication asks whether a finding survives when somebody tries again.
AI asks whether our system recognises scholarship because of the thinking behind it, or because it recognises the appearance of scholarship.
Neither test is particularly comfortable.
Perhaps that is why they are useful.
Replication failure should certainly change research design: better justified samples, adequately powered studies, preregistration where appropriate, transparent materials and data, clearer distinctions between confirmatory and exploratory work, replication and Registered Reports.
It should change analysis too. Less worship of statistical significance; more attention to effect sizes, uncertainty, robustness, assumptions and plausible alternative explanations.
And it should change teaching. Students should learn that science is not the production of significant results. They need scepticism, transparency and some tolerance for uncertainty.
But none of this will be enough if the reward system remains largely untouched.
If researchers are told to conduct slower, more careful science while career progression rewards volume, grant income and prestigious publications, we shouldn't be terribly surprised when behaviour follows the reward rather than the sermon.
If journals celebrate reproducibility while continuing to compete for novelty and attention, the same tension remains.
If universities tell us education is about intellectual development while increasingly valuing what can be placed conveniently into a dashboard, we have the same problem again.
And if academic publishing now responds to AI by defending the superiority of human scholarship without examining the system that helped produce the replication crisis in the first place, I think we will have missed something important.
AI did not create academic publishing's problems.
But it may be making some of them much harder to ignore.
Replication failure should have changed more than psychology's statistics. It should have changed what we reward, what we publish, what we promote and what we teach.
Most of all, perhaps it should have made us rather less certain that the things we have become good at measuring are necessarily the things that matter.
References
Coalition for Advancing Research Assessment. (2022). Agreement on Reforming Research Assessment.
Giray, L., Sevnarayan, K., Maphoto, K. B., & Wider, W. (2026). AI slop in academic publishing: History, characteristics, manifestations, causes, and mitigation strategies. Internet Reference Services Quarterly, 30(2), 221–244. https://doi.org/10.1080/10875301.2026.2637526
Giray, L., Sevnarayan, K., Maphoto, K. B., & Wider, W. (2026). Corrigendum to "AI Slop in Academic Publishing: History, Characteristics, Manifestations, Causes, and Mitigation Strategies". Internet Reference Services Quarterly. https://doi.org/10.1080/10875301.2026.2713779
Gartenberg, C., Hasan, S., Murray, A., & Pierce, L. (2026). More versus better: Artificial intelligence, incentives, and the emerging crisis in peer review. Organization Science. https://doi.org/10.1287/orsc.2026.ed.v37.n3
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, 0021. https://doi.org/10.1038/s41562-016-0021
Nosek, B. A., Alter, G., Banks, G. C., et al. (2015). Promoting an open research culture. Science, 348(6242), 1422–1425. https://doi.org/10.1126/science.aab2374
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Smaldino, P. E., & McElreath, R. (2016). The natural selection of bad science. Royal Society Open Science, 3(9), 160384. https://doi.org/10.1098/rsos.160384
Soderberg, C. K., Errington, T. M., Schiavone, S. R., Bottesini, J., Thorn, F. S., Vazire, S., Esterling, K. M., & Nosek, B. A. (2021). Initial evidence of research quality of registered reports compared with the standard publishing model. Nature Human Behaviour, 5, 990–997. https://doi.org/10.1038/s41562-021-01142-4
Thakkar, N., Yuksekgonul, M., Silberg, J., Garg, A., Peng, N., Sha, F., Yu, R., Vondrick, C., et al. (2026). A large-scale randomized study of large language model feedback in peer review. Nature Machine Intelligence, 8, 326–336. https://doi.org/10.1038/s42256-026-01188-x