Science is dead – long live science
On a publishing crisis that generative AI did not cause but has accelerated – and why it can still be a way out.
Science is in crisis. For once, I don’t mean trust in facts, which is under pressure in an age of disinformation and “alternative facts”. I mean a crisis in the engine room: in the way research is published, checked and evaluated. It is less visible, but it affects exactly what that trust is supposed to rest on.
Two announcements in one week
On 1 October 2026, arXiv, the most important preprint server for computer science, mathematics and physics, introduced a new rule: each person may now submit only two papers per calendar month and have at most three in moderation at the same time (Boboris 2026). Five days later, OpenAI published 722 mathematical manuscripts on GitHub, covering 372 families of results, generated by an internal, unreleased model (OpenAI 2026b; Schuler 2026).
Both announcements describe the same problem from two sides. Scientific texts can now be produced faster than people can read, check and contextualise them. The bottleneck in science is no longer the production of results but their verification.
The title of this post is deliberately provocative. I don’t believe science is dying. But I do believe that a particular model of science is reaching its limits: the model in which publications are counted, peer review is done unpaid on the side, and quality control relies on writing a paper being hard work. In my view, generative AI did not trigger this crisis; it exposed an imbalance that has existed for a long time. And that is exactly why it can also be part of the solution.
Before getting to the crisis, it helps to look at how scientific quality control works in the first place. Anyone who wants to publish a research result submits a manuscript to a journal or conference. The editors first check whether the paper fits the scope and meets basic standards. Some submissions are rejected at this stage, known as a desk reject.
The remaining manuscripts go to usually two or three experts from the same field, the reviewers or referees. They assess the research question, methods, data and conclusions and write a report with a recommendation: accept, minor or major revision, reject. Reviews are usually anonymous; often the reviewers do not know the authors’ names either (double blind). After one or more rounds of revision, the editors decide. Publication often takes months.
Three properties of this system matter for this post. First, reviewing is almost always unpaid. It counts as service to the research community and is done alongside research and teaching; Aczel, Szaszi, and Holcombe (2021) estimate around six hours per review. Second, reviewers are the same people who publish. Anyone who writes more therefore creates more reviewing work for others. Third, the process only scales with human attention, and that is limited.
Preprint servers such as arXiv work differently. Papers are made public before or alongside peer review so that results are available quickly. arXiv itself does not do peer review. Volunteer moderators only check whether a submission meets minimum scholarly standards, is of interest to the research community and falls into the right category (Boboris 2026). Of all things, it is this comparatively light check that is now reaching its limits.
How we got here
To understand today’s situation, a brief look back helps. For a long time, academic publishing was free for authors. What cost money was reading: libraries subscribed to journals, and those subscriptions became more and more expensive over decades.
At the same time, the market became concentrated. Based on 45 million documents, Larivière, Haustein, and Mongeon (2015) show that the five largest publishers published around 20 percent of all papers in 1973 and more than half in 2013. Digitisation did not democratise the dissemination of knowledge; at first, it mainly strengthened the large commercial players.
The business model is remarkably profitable. The Scientific, Technical & Medical division of RELX, which includes Elsevier, achieved an adjusted operating margin of 38.1 percent in 2025 (RELX 2026). A substantial part of the value is created for free: Aczel, Szaszi, and Holcombe (2021) estimate that reviewers worldwide spent more than 100 million hours on peer review in 2020. For the US alone, that corresponds to more than 1.5 billion US dollars in salary value.
Open access shifted the problem
The open access movement was the right answer to closed science. In Germany, Projekt DEAL negotiated nationwide agreements with the major publishers. The first agreement, with Wiley in 2019, set a publish-and-read fee of €2,750 per article; before that, publishing itself had been free for authors (Dobusch 2019). The Elsevier agreement followed in 2023 at €2,550 per article (open-access.network 2023).
This reversed the logic. Reading is free, publishing costs money. For publishers, this means revenue no longer depends on the number of readers but on the number of published articles. RELX itself describes growth in primary research as volume-driven and points to strongly rising submissions, especially in the pay-to-publish segment (RELX 2025).
This logic produced strange outgrowths long before generative AI. So-called predatory publishers sell publications without meaningful peer review and organise conferences where the registration fee effectively guarantees publication. Against the best-known of them, the OMICS Group, a US federal court imposed more than 50 million US dollars in 2019 in a case brought by the Federal Trade Commission, among other things for faking peer review (Federal Trade Commission 2019). And as early as 2014, Springer and IEEE had to retract more than 120 conference papers that the program SCIgen had generated as meaningless text. They were not caught in peer review but only years later (Van Noorden 2014). Predatory publishers are a fringe phenomenon. But they show in extreme form where an incentive leads that also affects reputable publishers: if you earn money from every publication, you have little interest in fewer things being published.
Hanson et al. (2024) described this development as the strain on scientific publishing. The number of articles indexed in Scopus and Web of Science in 2022 was about 47 percent higher than in 2016, while the number of active scientists has barely grown. The load per person for writing, reviewing and editing is therefore rising. That was the state of affairs before generative AI became widely available.
The flood: arXiv pulls the emergency brake
Generative AI is therefore hitting a system that was already under pressure. This is most visible at arXiv, because the platform publishes its numbers. In September 2016 it received 9,869 submissions, in September 2024 20,569 and in September 2026 40,363 – a new record. The number has doubled in just two years; in the cs.AI category it has grown more than sixfold (Boboris 2026).
arXiv stresses that AI is explicitly allowed as a tool, as long as its use is disclosed and the paper makes a scholarly contribution. Many of the new submissions, however, do not meet these requirements. Moderators are seeing more thin papers with a narrow focus, more “salami-sliced” papers and significantly more dense, AI-written text. In September 2026 alone, almost 9,000 support tickets landed on a team made up mostly of volunteers (Boboris 2026).
The limit is only the latest step so far. Since October 2025, arXiv has accepted review articles and position papers in computer science only if they have already been peer reviewed and accepted by a journal or conference; hundreds of such texts were coming in every month, many little more than annotated bibliographies (Boboris 2025). Since January 2026, an institutional email address is no longer enough for new authors. They need a previously accepted arXiv paper or a recommendation, a so-called endorsement, from an established author (arXiv 2025). And since May 2026, authors face a one-year ban if their submissions show clear signs of unchecked AI output, such as fabricated references or leftover chatbot comments (Grove 2026).
arXiv is not an isolated case. bioRxiv and medRxiv, the preprint servers for biology and medicine, were already rejecting more than ten formulaic, possibly AI-generated manuscripts a day in 2025 (Watson 2025; summarised in Price 2025). At the smaller Engineering Archive (engrXiv), monthly submissions rose from an average of 60 to almost 200 (engrXiv 2025).
What stands out is how arXiv counts towards the new limit: every submission counts, not just every accepted paper. A rejected text therefore also uses up one of the two monthly slots, because checking it also costs moderation time. The reasoning is that a relatively small share of authors submits a large number of low-quality papers, delaying everyone else’s work by days or weeks (Boboris 2026). The limit is controversial: critics argue that a blanket cap does not distinguish quality from quantity and hits small groups harder than large ones (Times Higher Education 2026).
AI is also arriving on the reviewing side. Pangram Labs, a provider of AI detectors, classified around 21 percent of the reviews for the AI conference ICLR 2026, 15,899 in total, as fully AI-generated; more than half showed at least some signs of AI use (Pangram Labs 2025; Naddaf 2025). When AI-written papers are assessed by AI-written reviews, it is fair to ask who is actually still checking.
At the same time, a market is emerging that monetises the overload. In February 2026 alone, Pottarath, Halffman, and Horbach (2026) found more than 1,000 LinkedIn ads for “freelance” reviewers, mostly posted by service providers that take over parts of peer review for publishers. The reviewers interviewed report €30 to €40 per review and deadlines of usually two to three days. As the authors put it, reviewing thus changes from a service to the research community into an individual, fast transaction. Research integrity experts warn that such quick reviews could be written with AI without anyone noticing and used by paper mills to give fabricated papers a veneer of respectability (Naddaf 2026).
Publications as currency
Behind the flood lies an incentive system that is older than any AI. In science, publications are the currency in which performance is measured. They decide doctorates, positions, grants and professorships. Anyone pursuing an academic career has to show enough visible work within a limited time, usually on fixed-term contracts.
The greatest pressure is felt not by established researchers but by younger ones. In a survey at four Amsterdam institutions, postdocs and assistant professors reported the highest publication pressure; PhD candidates were the most likely to lack the resources to cushion it (Haven et al. 2019). And the pressure is growing: in the Researcher of the Future report, 68 percent of more than 3,200 respondents said they were under more pressure to publish than two to three years earlier. Only 45 percent felt they had enough time for research (Elsevier 2025).
The open access model makes things worse for exactly this group. Where publishing costs money, the visibility of a paper also depends on whether someone can pay the fee. In an AAAS survey of 422 US researchers, just over half of those who had already paid publication fees found it difficult or very difficult to find the funds. More than three quarters had given up materials, equipment or tools to do so (AAAS 2022). Researchers without their own project budget, often PhD candidates and postdocs, depend on institutional agreements such as DEAL or on their group’s budget.
Against this background, it seems plausible that publication pressure also lowers the threshold for using AI. As far as I know, this link has not yet been demonstrated directly, but the evidence points in that direction. Higher publication pressure is associated, albeit weakly, with a greater willingness to violate ethical standards (Paruzel-Czachura, Baran, and Spendel 2021). Researchers with less than three years of experience use AI for research and editing text more often than experienced ones, 48 versus 34 percent (Springer Nature 2026). And a simulation by Jiang (2025) shows how strongly a tenure system based on publication counts would reward early AI adopters.
This is not a criticism of young researchers but a rational response to a system that rewards quantity. Anyone who does without AI while others use it risks falling behind in the competition for a small number of permanent positions. That is why the problem cannot be solved by rules for authors alone, only by changing how research is evaluated.
Honest research gets faster too
It would be convenient to blame the problem solely on paper mills and dishonest authors. But that falls short. Even researchers who work ethically, disclose their AI use and check every result themselves are getting faster.
I notice this in my own work. Literature searches, data preparation, analysis code and first drafts are much faster with AI support than two years ago. The ideas and the responsibility for the results remain mine. But the path from an idea to a submittable manuscript has become shorter.
For the system as a whole, the arithmetic is simple. If every researcher legitimately completes 30 or 50 percent more manuscripts, the number of reviews needed rises by the same amount. But the reviewers are the same people who are currently writing more themselves. The time it takes to do a careful review has not decreased; if anything it has increased, because polished text requires closer scrutiny. Capacity for checking does not grow with capacity for production.
From literature reviews to mathematical proofs
For a long time, the problem seemed limited to certain types of text. Review articles, literature searches and incremental empirical papers were easy to speed up with language models. Mathematical proofs, on the other hand, were seen as an area where AI could at most assist.
That changed fundamentally this autumn. After ten results in August and a claimed contribution on the Navier–Stokes equations in September, OpenAI presented 372 families of results on 6 October. Each is supposed to solve or substantially advance an open mathematical question, including a claimed proof of the Unique Games Conjecture (Schuler 2026). According to OpenAI, the model was given around 4,000 problems; on average, a result required computing time equivalent to about three hours of ChatGPT Pro reasoning (OpenAI 2026b, 2026a).
Lean comes up again and again in the debate about AI proofs. It is an open-source programming language and a so-called proof assistant, now developed by the non-profit Lean FRO (Lean FRO 2026). Mathematical statements and proofs are written down in it so precisely that a computer can check every single step for logical correctness. The foundation is the community-maintained library Mathlib, which now contains around 290,000 formalised theorems (Lean Community 2026). If Lean accepts a proof, nobody has to rely on a reviewer’s judgement for the correctness of the individual steps.
For my argument, the question of how many of the OpenAI results will ultimately hold up matters less. What matters more is where the checking work ends up. According to the repository, around 42 percent of the main results are formalised in Lean, i.e. checked step by step by a computer. For the remaining results, OpenAI itself concedes that they may contain errors (OpenAI 2026a). Three manuscripts were withdrawn after a sign error only one day after publication, and 14 more were revised (Schuler 2026).
Even the formally checked proofs are not automatically understood. One day after publication, complexity theorist Aaronson (2026) wrote that apparently no human had yet understood nearly any of these proofs. On his blog, he also quotes Dana Moshkovitz, who worked on the Unique Games Conjecture for years: she described the proof as so badly written that she could only read it with the help of AI. According to an OpenAI spokesperson, the company’s own mathematicians have not yet worked through many of the results either (Schuler 2026).
This is where the asymmetry becomes especially clear. Generating more than 700 manuscripts cost computing time. Checking, contextualising and explaining them costs the lifetime of experts who were neither asked nor paid to do so. Lean helps because it mechanically guarantees the logical correctness of a proof. Whether the formal statement matches the intended conjecture, whether the result is interesting and what it means for the field still has to be judged by a human. An advisory group at the Institute for Advanced Study had therefore called in advance for the model, prompts, compute and a summary of the solution path to be disclosed for every result. OpenAI published averages and ten summaries, but no prompts (Schuler 2026).
How deep the unease in mathematics runs is shown by the statement A Severe Misalignment of AI in Mathematics, published on 11 September 2026, a few days after the Navier–Stokes announcement. It has now been signed by 28 Fields medallists, including Terence Tao, Peter Scholze and Maryna Viazovska. They consider it harmful that AI companies use mathematical problems as benchmarks: solving problems, they argue, is only a tool and proxy for the actual goal, conceptual understanding. Mass-produced results could destroy rather than prepare the ground for new ideas. Without mathematicians who absorb AI-generated ideas and pass them on, the human chain of transmission in the field would also be lost. Rushed announcements leave no time for proper write-up and citation and raise serious questions about authorship and plagiarism (Avila et al. 2026). The signatories do not call for a ban. Whether AI helps or harms mathematics, they say, depends on human choices; the research community, AI companies and society need to act urgently.
The same pattern in companies: workslop
One could dismiss all this as a niche academic problem. I think the opposite is true. Science is simply the area where the problem is measured best.
In companies, the same mechanism is at work, just without public statistics. Wherever AI is used, more output is produced: more presentations, more reports, more emails, more code, more concept papers. Checking capacity, on the other hand, stays largely the same. It lies with managers, colleagues and subject-matter experts who have to read, understand and take responsibility for the result.
Niederhoffer et al. (2025) coined the term workslop for this in the Harvard Business Review: AI-generated work that looks like good work but lacks the substance to meaningfully advance a task. In their survey of US employees, around 40 percent said they had received such output in the past month; each case took almost two hours of rework on average. Extrapolated to respondents’ salaries, Axios (2025) put the invisible cost at around 186 US dollars per person per month.
The core of the problem is the same as at arXiv. Workslop shifts cognitive work from the person who produces something to the person who receives it. Anyone who passes on an AI-generated report without checking it saves an hour and costs the recipient two. As in science: producing has become cheap, checking has not. A company that measures AI mainly by the amount of output it produces is therefore unintentionally recreating the situation arXiv is currently capitulating to.
Long live science: how AI could save science
Despite all this, I am optimistic. The crisis reveals that quality control in science rested on an assumption that no longer holds: that the effort of writing acts as a natural filter. When that filter disappears, we have to design checking deliberately. That is an opportunity.
Checking, not just producing. The most obvious consequence is to use AI where the bottleneck is. Tools can detect duplicates and heavy overlap, check references automatically against databases, test statistical reporting for consistency and reproduce code. This does not replace expert review, but it can relieve moderators and reviewers of routine work. Many comments on the arXiv rule call for exactly this: AI not only on the authors’ side but also on the side of those doing the checking (Boboris 2026).
Use formal verification where possible, without overrating it. In mathematics, Lean can take over part of the checking. As we saw with the OpenAI results, however, it only checks what has been formalised, and only whether a proof matches the formally written statement. Formalisation is laborious but is getting faster with AI: for Fermat’s Last Theorem, Kevin Buzzard launched a project in 2024 planned to take years; in September 2026, Claude formalised the proof in eleven days, albeit with around 13 million lines of Lean code, a good five times the size of all of Mathlib. Anthropic itself stresses that formal proofs should not replace human-readable explanations (Anthropic 2026). Lean itself is not free of bugs either: in July 2026 it accepted an AI-assisted “disproof” of the Collatz conjecture that in fact exploited a bug in the program’s kernel. It was fixed within an hour (Moura 2026). Above all, though, a formal check does not replace understanding; more on that below.
The empirical sciences have nothing comparable. No machine can prove that an experiment was carried out properly or that a finding generalises. But there are building blocks for routine checking: executable analysis code, open data, preregistered analysis plans and automated consistency checks. Even simple tools find a lot. Using statcheck, Nuijten et al. (2016) found at least one p-value that did not match the reported test statistic in about half of the psychology articles examined; in about one in eight articles, the discrepancy could change the conclusion. Such tools check consistency, not truth. But they relieve people exactly where checking is mechanical.
New data and new questions become more valuable. AI is particularly good at recombining what is already known: summarising literature or analysing public datasets according to the same pattern over and over. Suchak et al. (2025) show what this looks like using the freely available US health database NHANES. Between 2014 and 2021, an average of four studies per year linked a single factor to a disease; in 2024 there were already 190 by early October. Many of them do not correct for multiple testing or select time periods without justification. What AI cannot generate, by contrast, are new measurements, experiments and observations. Research that produces new data is therefore likely to gain in value.
This does not mean that reanalysing existing data becomes worthless. After the replication crisis, careful replications are more important than ever; the line runs between formulaic and thoughtful, not between old and new data. And it is not only new data that remains scarce, but also new questions: the OpenAI proofs show that AI can produce results without new data, but not automatically understanding. This does favour well-funded groups, though, because collecting new data is expensive. And the more valuable data becomes, the more important its traceable provenance, because plausible-looking data is also easier to fabricate with AI.
This implies a shift in scientific work itself. If AI solves problems faster, finding the right problems and formulating them precisely becomes more important: Which question is worth asking, which data is missing, what would even be an interesting result? This applies far beyond science. In everyday work with AI, too, the quality of the question increasingly determines the quality of the answer. The Fields medallists’ statement does point to a catch, though: the ability to ask new questions has so far developed precisely through the laborious work of tackling problems oneself. If AI takes over this training, that very ability risks withering away (Avila et al. 2026). I made a similar argument about teaching programming: students need to learn to code precisely because AI can write code. For training young researchers, this is one of the most important open questions.
Change the incentives. The flood is also a consequence of publish or perish. As long as careers depend on the number of publications, it is rational to publish more with AI. arXiv explicitly says that authors should submit only their best work. ICML 2026, one of the major AI conferences, goes further and for the first time includes in its decisions rankings in which authors rank their own submissions by quality. In an experiment at ICML 2023, such self-assessments predicted later citations better than the reviews did (Su et al. 2025). Hiring and funding processes that evaluate only a small number of selected papers would noticeably reduce the pressure for quantity.
Make reviewing work visible and pay for it. If peer review is the real bottleneck, it should be treated as such. The journal Biology Open, for example, now pays reviewers for thorough reviews delivered within four working days (Naddaf 2026). What matters is that this happens transparently and does not end up in an opaque market for quick reviews.
Separate results from explanations. Finally, the OpenAI release shows that a correct result alone is not yet scientific progress. Knowledge only emerges when people can understand, contextualise and pass on a result. In future, the real scientific achievement may more often lie in checking, explaining and connecting machine-generated results to existing knowledge. Cryptographer Dakshita Khurana, who worked for years on one of the problems now claimed, puts it roughly like this, according to Schuler (2026): even if someone else has reached the summit by a different route, from up there you can see new mountains that you can now climb with students, colleagues and machines.
What companies can learn from this
Some fairly concrete lessons for companies can be drawn from the crisis in science:
- Checking capacity is the scarce resource. Anyone introducing AI should ask not only how much faster results are produced, but who checks them and how much time they have for it. Review capacity should be planned just like budget or compute.
- Responsibility lies with the person who passes it on. arXiv counts submissions against the person who clicks “submit”. Translated: whoever passes on an AI-assisted piece of work is responsible for its content, checks it themselves first and discloses where AI was involved. This is not a gesture of mistrust; it helps the recipient estimate how much checking is needed.
- Measure outcomes, not output. The number of documents, slides or lines of code produced are poor measures of success. More relevant are decision quality, lead time to a result that is actually used, and error rates.
- Automate checking first, then scale production. Tests, evaluations, checklists and plausibility checks are the operational counterpart to Lean. As long as they are missing, limits are legitimate too: a cap on items in progress at the same time, like arXiv’s three active submissions, forces people to select the best rather than pass everything on.
- Cultivate your own data and knowledge. Generic analyses and texts become interchangeable because everyone can produce them with the same models. What makes the difference is your own data, your own process knowledge and the ability to ask the right new questions.
Conclusion
Science is currently experiencing what many organisations still have ahead of them. Generative AI is dramatically lowering the cost of producing text, code and even proofs. The cost of checking is not falling at the same rate. A system whose quality control quietly relied on the effort of writing is being thrown off balance.
arXiv’s emergency brake is therefore understandable, but it is not a solution. A limit reduces quantity but does not distinguish between good and bad work. In the long run, what will matter is scaling checking itself and leaving the responsibility for it with those who produce something.
In that sense, the title is less pessimistic than it sounds. A particular way of organising science is coming to an end. What comes next can be better – if we use AI not just for writing but above all for checking. The same is true for companies.