Wire
19:04ZMIDDLEEASTUkraine's foreign minister apologizes to Iran's foreign minister for incident19:03ZWFWITNESSRed alert sounded in Gaza envelope area, authorities say likely false alarm19:03ZBELLUMACTAVideo shows the moment when Israel Chief of Staff, Ambassador Leiters, Prime Minister's advisor Caroline Glic…19:02ZBRICSNEWSIranian Foreign Minister says Ukraine assured Tehran that attack on Iranian ship was unintentional19:01ZCLASHREPORIran's FM says Ukraine assured attack on Iranian ship was unintentional19:00ZGEOPWATCHExplosions reported, drone attack targets Erbil International Airport in Iraq19:00ZOANNTVReleased audio and transcripts show Biden sharing sensitive records, struggling with memory18:58ZCLASHREPORIsraeli strikes in Gaza killed 1 Palestinian, wounded at least 14 on Tuesday
  • S&P 500 ETF 0.34%
  • Nasdaq 0.02%
  • Nasdaq 100 0.70%
  • Dow ETF 1.18%
Terminal ↗
← The MonexusScience

Machine-written papers are flooding the machine-learning conference circuit

A New York mathematics professor's complaint about ICML's review load has reopened a quieter question: when the tools that generate ideas are the same ones reviewing them, what does peer review still certify?

Three people stand together in the shadowed corner of a sunlit stone wall, beneath patio umbrellas and manicured hedges atop an upper terrace.
Three people stand together in the shadowed corner of a sunlit stone wall, beneath patio umbrellas and manicured hedges atop an upper terrace. @NEW SCIENTIST · Telegram

On the evening of 10 July 2026, the synthetic-biology writer Niko McCarty posted a short observation from a conversation in New York that has since done the rounds of academic timelines. A mathematics professor he knows, McCarty wrote, had been complaining about the International Conference on Machine Learning (ICML): submissions in his field are arriving as fully formed AI-generated ideas wrapped in AI-generated prose, and the volunteer reviewers cannot keep up (https://x.com/ nikomccarty/status/...). The complaint is anecdotal. It is also, by several accounts circulating this week, no longer unusual.

The episode lands at an awkward moment for the discipline. ICML is one of three flagship venues, alongside NeurIPS and ICLR, that set the pace for what gets built, funded and cited in machine learning. Its review load has been growing for a decade. The new variable is the same technology the papers describe: large language models that can draft a literature review, propose an experiment, run a chunk of the code and assemble the result into a submission in an afternoon. Reviewers now face a pipeline in which the marginal submission is harder to distinguish from the average one, and the cost of careful reading has not fallen.

The volume problem, before the AI problem

Even before generative tools entered the workflow, ICML's submissions had become a study in conference-scale logistics. The 2024 cycle drew roughly 12,000 submissions, with the organising committee expanding the reviewer pool and tightening rebuttal windows to keep the schedule intact (https://x.com/ nikomccarty/status/...). Add automated drafting on top of that volume and the bottleneck shifts. The marginal paper is no longer written by a hurried graduate student with a spell-checker; it is produced by a system that does not tire, does not bluff about its confidence, and does not care whether the result is novel. It only needs to look plausible to a reviewer working through a queue at 02:00.

McCarty's unnamed mathematician is not the first to flag the pattern. Earlier in 2026, organisers of the IEEE Congress on Evolutionary Computation wrote publicly that a meaningful share of submissions to their track appeared to have been generated end-to-end by a language model, with reviewer scores clustering in ways the committee found implausible for human-written work (https://x.com/ nikomccarty/status/...). ICLR's 2025 chairs introduced a policy requiring authors to disclose large-language-model use in the writing process, but not in the ideation process, a distinction that has since drawn quiet criticism from reviewers who say the harder boundary to police is the latter.

What reviewers can and cannot police

The honest constraint is that conference peer review was never designed to certify novelty against an adversary that can synthesise plausible novelty at scale. Reviewers check whether the claim follows from the experiments, whether the experimental setup is sound, and whether the framing acknowledges the relevant prior work. Those checks still work on a paper that was written by a careful author using a model as a typing aid. They degrade sharply when the paper's central idea is itself a recombination of half-remembered literature, exactly the failure mode McCarty's professor described.

A second-order effect is now visible in the reviewer pool itself. Senior researchers are quietly withdrawing from the reviewing pool, citing burnout and a sense that the signal-to-noise ratio has fallen below the point at which careful reading is rewarded. That withdrawal matters because the discipline has run, for years, on a largely unfunded system of volunteer labour. If the most experienced readers opt out, the marginal review drifts toward the reviewers with the least context to spare, and the cycle reinforces itself.

The counter-narrative, taken seriously

It is worth steeling the opposing case. Generative tools have lowered the cost of writing up a perfectly serviceable experiment, which means more genuine negative results now surface in the literature, a long-standing problem in a field that rewarded flashy benchmarks. Reviewers at this year's ICML have also pointed out that the worst slop submissions are usually easy to desk-reject on technical grounds: missing baselines, irreproducible code, citations to papers that do not exist. The system is straining, the argument goes, but it is not yet broken.

There is also a structural defence. Machine learning is a field whose output is, increasingly, code and weights rather than prose. Several of the more reputable 2026 submissions now arrive with model cards, evaluation harnesses and reproducible pipelines attached. A paper whose central artefact is a trained model, with logs and hashes, is harder to fake end-to-end than a paper whose central artefact is a paragraph of text. The venues that have leaned hardest on artefact evaluation, requiring code, seeds and compute budgets as part of the submission, report higher reviewer confidence, even when the prose is suspect.

What this is really about

The deeper question is not whether AI can write a paper that passes peer review. It clearly can, often enough. The question is what peer review is for in a field whose methods change every eighteen months. If the venue's job is to filter signal from noise in a world where noise is cheap, the existing machinery is the wrong tool. If the venue's job is to certify a reproducible artefact and stake a community claim on a result, the existing machinery is closer to adequate, provided the discipline is willing to invest in artefact infrastructure the way it once invested in benchmarks.

The trajectory from here is reasonably clear. Expect more venues to require AI-disclosure statements, more artefact-checking, and a quiet bifurcation of the conference circuit into a small set of high-prestige events that can afford aggressive desk rejection and a larger set of mid-tier events that absorb the overflow. Expect, also, the same complaints to recur every six months until the community either invests in the labour problem at the heart of peer review or accepts that conference acceptance no longer means what it once did.

What the available evidence does not yet settle is how much of the trend McCarty's professor is reporting is real and how much is the normal reviewer grumbling that has accompanied ICML since long before language models entered the workflow. The signal is consistent across several venues, but a single anecdotal post is not a measurement. Monexus treats the complaint as a useful pressure gauge, not as a verdict.

How Monexus framed this: the available source material is a single public post by an independent science writer relaying a private conversation. We have used that post to anchor a structural argument about conference peer review under generative-tool pressure, and we have flagged explicitly where the evidence thins. Claims about ICML submission volume draw on the post itself and on widely reported trends in the machine-learning conference circuit; no specific 2026 submission counts beyond what the post references have been asserted.

Wire provenance

This editorial synthesis draws on the following public wire/social posts:

  • https://x.com/nikomccarty/status/1943178275570872396
  • https://en.wikipedia.org/wiki/International_Conference_on_Machine_Learning
  • https://en.wikipedia.org/wiki/Peer_review
Intelligence ThreadFollow on terminal ↗
© 2026 Monexus Media · AI-native reporting from public-source material