Key Takeaways
- A careful reviewer's most valuable move is demanding the one ablation authors would never volunteer on their own because it might undercut their own result.
- Much of review's real value happens invisibly between submission and camera-ready, when overclaims get walked back before most readers see the early draft.
- Review cannot verify that reported numbers are real or confirm a result generalizes beyond its test distribution, so acceptance should read as a modest signal, not a verdict.
It has become fashionable, and not entirely unfair, to treat peer review in AI as a formality that happens after the fact. The paper is already on a preprint server, already read by everyone who matters, already cited in three follow-up projects and maybe already shaping a product roadmap, and months later a venue attaches an accept decision that changes essentially nothing about how the work was received. Given that sequence, it's a reasonable question: why defend a process that arrives this late to its own party?
Because "everyone already read it" and "everyone already checked it" are different claims, and quietly treating them as the same one is exactly the mistake that lets weak papers travel further and faster than they should. Review does something specific that a viral thread, an excited launch post, and a flurry of citations cannot do on their own — it puts the paper in front of people with no stake in the outcome and an explicit obligation to write down what's wrong with it. That's worth defending carefully, without pretending the process doesn't have real and well-documented problems.
The Complaints Are Mostly Fair
Start with the case against, because most of it holds up. Submission volume has grown far faster than the pool of qualified reviewers, which means a given paper's fate increasingly rests on two or three people, at least one of whom is likely reviewing outside their exact subfield and pattern-matching against surface features rather than engaging with the details. Rebuttal periods reward polished pushback and confident tone as much as they reward substance. The whole cycle takes months in a field that now iterates in weeks, which means the formal decision frequently lands well after the finding has already been absorbed, argued over, and built upon by everyone who was going to read it anyway. None of these complaints are exaggerated. They describe the process as it actually operates at most major venues, most of the time. Anyone who has sat on a program committee during a submission crunch has watched a paper get a rushed, superficial read simply because the reviewer assigned to it was already carrying four other submissions with the same deadline, and no amount of process design fully fixes a shortage of qualified attention.
What a Careful Reviewer Actually Catches
And yet a genuinely careful review does something distinctive that the informal, preprint-first reading of a paper structurally tends to skip. A good reviewer checks for leakage between training and test data, a mistake that's easy for authors to miss precisely because they're too close to their own pipeline to see it. A good reviewer asks for the specific ablation that would reveal whether a claimed component is actually doing anything, or whether the result would survive with that component quietly removed — the exact experiment authors are least likely to volunteer on their own, because it might undercut the story they're telling. A good reviewer checks whether the statistical comparison accounts for run-to-run variance, whether the claims in the abstract are actually supported by the numbers in the tables, and whether closely related prior work got cited in a way that honestly represents how novel the contribution really is.
None of that is glamorous. It's also precisely the kind of scrutiny that a fast, excited, informal read of a preprint tends not to apply, because informal readers are looking for what's interesting, not for the specific hole in the argument. Those are different reading modes, and a field that only ever reads in the first one is a field that gets fooled more often than it should.
The Adversarial Reader Problem
Authors are structurally the worst-positioned people to find the holes in their own paper, and this has almost nothing to do with dishonesty. By the time a paper is ready to submit, its authors have spent months getting intellectually and emotionally invested in the story of why their method works. They've stopped noticing the alternative explanations they ruled out early and never revisited. They've absorbed a framing that flatters the result, not because they're trying to deceive anyone, but because that's what sustained, motivated attention to your own work does to your perspective on it.
A reviewer with no investment in whether the paper succeeds, whose entire job for the next few weeks is to find the hole, is a genuinely different kind of check than "we reread it ourselves one more time" or "the response online has been positive." Public reaction rewards excitement, confirmation, and novelty. It does not reward the boring, specific question of whether the ablation table in section five actually supports the claim made in paragraph four of the introduction. Somebody has to be assigned that question, with no upside for answering it kindly, or it mostly doesn't get asked at all.
Where Review Runs Out of Road
None of this should be mistaken for review being sufficient, and it's worth being precise about exactly where it stops working. A reviewer cannot verify that reported numbers are real. Nobody is rerunning your code on your cluster with your exact data from a PDF over a three-week review window, which means convenient rounding, cherry-picked seeds, or outright fabrication are all nearly invisible to the process regardless of how careful the reviewer is. Review cannot tell you whether a result generalizes past the specific distribution it was tested on. It cannot move at the speed the field currently operates at, which means it structurally cannot be the primary gate for decisions practitioners need to make in real time, long before any formal decision comes back. And it doesn't fully escape its own social dynamics — reviewers have documented tendencies to extend more benefit of the doubt to submissions from well-known labs and to apply sharper scrutiny to unfamiliar names, which is exactly the kind of bias an adversarial process is supposed to guard against and only partially does.
The Paper You Cite Isn't the Paper That Was Submitted
There's a quieter contribution review makes that gets almost no credit, because it's invisible unless you go looking for it: a meaningful share of what review actually does happens between a paper's initial submission and its final camera-ready version, regardless of whether the accept-or-reject outcome was ever in real doubt. A reviewer flags an overclaim in the abstract that doesn't match the evidence, and the final sentence gets walked back. A reviewer asks why a specific competing method wasn't included in the comparison table, and it shows up in the revision. A reviewer points out that a limitation deserves its own paragraph instead of a single hedged clause, and readers a year later benefit from that disclosure without ever knowing a reviewer is the reason it exists.
None of that shows up in the binary accept-or-reject number people cite when they argue review doesn't matter, because the version of the paper that critique is usually aimed at is the polished final one, after review has already done its quiet editing. Preprint-first publishing has made this contribution easier to miss, since the early version circulating on social platforms may already have been through a round or two of exactly this kind of pressure by the time most casual readers encounter it, or may represent a version from before review that a later revision meaningfully improved on. Either way, the credit rarely follows the actual work.
Acceptance Is a Signal, Not a Verdict
Put those two halves together and a workable posture starts to emerge. Treat "peer reviewed" as modest, genuinely useful evidence that a few qualified strangers looked hard for specific holes and didn't find anything fatal enough to block publication — not as a certificate that the paper is true, and not as a substitute for reading the results and limitations sections yourself. And resist the opposite instinct too: don't let "just a preprint" become an automatic dismissal, since a meaningful share of the field's most important recent work has never gone through formal review at all and likely never will, given how the field's publishing habits have shifted. The right posture is neither reflexive trust in a venue's stamp nor reflexive suspicion of unreviewed work, but the same close reading any serious claim deserves regardless of which path it took to reach you.
That posture is harder to hold onto than it sounds, because both shortcuts are genuinely tempting. Trusting the stamp saves time in a field that moves too fast to closely read everything, and dismissing anything unreviewed feels like a defensible rule you can apply without having to make a judgment call on every individual paper. Neither shortcut is actually free. The first one occasionally lets a fatally flawed but well-credentialed paper coast into your understanding of the field unchallenged. The second one guarantees you'll be late to some of the most consequential work of the next few years, since a growing share of it simply won't take the traditional path anymore.
The honest version of this argument isn't "trust peer review" and it isn't "ignore peer review." It's that review is one imperfect, adversarial pass performed by people with nothing to gain from agreeing with the authors, and in a field this fast-moving and this financially motivated to overclaim, an imperfect adversarial pass is still doing real work that nothing else in the current system reliably does. Losing it wouldn't make the field's claims more true, and it wouldn't make anything move faster in any way that matters. It would just remove one of the few remaining moments where someone whose actual job is to find the hole in your argument is looking for it before the rest of us have to.