Key Takeaways
- The first filter is not how impressive a result sounds, but whether it would change an actual decision a practitioner faces this week.
- Open weights, released code, and disclosed evaluation details function as a coverage signal, not just good scientific practice, because they let an editor verify a claim before repeating it.
- A meaningful share of coverage is deliberately skeptical, explaining why an already-viral claim doesn't hold up rather than celebrating a new result.
Hundreds of new machine learning papers land on preprint servers every single day, and that's before counting workshop submissions, technical reports posted straight to a company blog, and GitHub repositories with a README that reads suspiciously like an abstract. Almost none of them are going to change what you do on Monday morning. A small number of them will change what an entire industry does for the next year. The problem for anyone trying to cover this field honestly is that on the day both kinds of paper get published, they look almost identical: a PDF, an abstract, a few charts, a confident claim.
We read this pile every day so that what reaches you has already survived a filter, and we think it's worth explaining exactly what that filter is, because "we decided it was interesting" isn't an editorial standard, it's an excuse dressed up as one. Here's what we actually check before a paper turns into a piece, and just as importantly, what gets it filtered back out.
The Real Question: Would This Change a Decision on Monday Morning
Our first filter has nothing to do with how impressive a result sounds and everything to do with whether it changes an actual decision somebody is facing this week. Does it change which model a team should fine-tune for a specific task? Does it change whether a particular prompting technique deserves the engineering time to try? Does it make a claimed safety concern real enough to budget for, or a claimed capability close enough to production-ready to prototype against? Picture a mid-size engineering team trying to decide whether to build an internal tool around a newly published technique versus waiting a quarter for it to mature — that's the reader we're actually writing for on most days, more often than the specialist who already tracks the literature closely.
A paper can be technically elegant and mathematically clean and still fail this test completely, because nobody downstream is actually deciding anything differently once they've read it. That's not a knock on the research, and it doesn't mean the paper is bad; plenty of foundational work looks purely theoretical for years before it becomes the thing an entire product category is quietly built on. It just tells us where a given piece belongs today, and today it usually isn't the top of the queue.
Novelty and Newsworthy Are Different Axes
"First paper to do X" is necessary information but rarely sufficient on its own, because plenty of firsts are firsts specifically because nobody else wanted to do that particular thing. Meanwhile a technique can be old news inside a specific research community and still be genuinely new information for everyone else the moment it gets cheap enough to run at scale, or the moment a lab with real distribution actually ships it, or the moment it quietly resolves an argument that's been running in public for months. We'll sometimes cover an idea that specialists have known about for a year, the week it becomes practically true for the rest of the industry, because timing and reach are their own kind of relevance and academic priority isn't the only axis that matters to a reader deciding what to build. The reverse happens too: a paper can be the first to formally demonstrate something the field has been informally assuming for a while, and the fact that it isn't a surprise to specialists doesn't make it any less worth writing up, because most of our readers aren't specialists and the formal confirmation is genuinely new information to them even when it isn't to everyone.
How a Paper Actually Moves From the Pile to the Page
The mechanics are less mysterious than they probably sound. Most papers get judged and set aside within a couple of minutes, on the abstract alone, using the same reading order we'd recommend to anyone: skip to the actual claim, check whether it's specific, check what it's being compared against. The ones that survive that pass get a second look at the results table and the limitations section, because that's usually enough to tell whether a promising-sounding abstract is backed by a real margin or a rounding error dressed up in confident language. Only a small fraction make it past that stage to a full read of the method and the appendix, and fewer still get cross-checked against what independent voices in the field are already saying about it, since a paper that looks strong in isolation sometimes looks very different once you see how specialists who aren't its authors are reacting to it.
What survives all of that isn't automatically framed as a triumph. Some pieces exist specifically to explain why a result matters and what it changes. Others exist to explain why a widely shared claim deserves more skepticism than it's getting, which is its own kind of coverage decision described further down. Either way, the piece that reaches you reflects a specific judgment call made under time pressure, by people who read a lot of these and occasionally still get the call wrong, which is exactly why the standard has to be explicit enough to argue with.
That last point matters more than it might sound. A stated standard is only useful if it's specific enough that someone else could look at a call we made and tell us we got it wrong, and specific enough that we could tell ourselves the same thing on a slow re-read a week later. Vague standards like "we cover what's important" don't clear that bar, because they can justify almost any decision after the fact. The version above is meant to be argued with, not just trusted.
Reproducibility Is a Coverage Signal, Not Just a Virtue
If a paper ships with open weights, released code, and a fully disclosed evaluation setup, that's not just good scientific practice, it's information we use directly in deciding how to cover it. It means we can actually check a claim ourselves before repeating it as more than "the authors say," and it means the authors were confident enough in the result to let outsiders try to break it. A closed, extraordinary claim without any of that gets treated with real caution, and if we cover it at all, we cover it explicitly as a claim under scrutiny rather than as settled fact. The same red flags that matter when any of us personally reads a paper apply doubly hard here, since the cost of getting it wrong isn't just our own understanding, it's whatever a reader does next based on how we framed it.
What We Actively Decline to Cover as Breakthroughs
A few categories get filtered out by default, and it's worth being specific about them rather than vague. A marginal benchmark bump purchased with dramatically more compute and no transferable insight about why it worked doesn't clear the bar, no matter how clean the chart looks. A known technique repackaged under new branding, with the lineage to prior work quietly left out of the framing, doesn't either. A single paper making a claim that contradicts a wide, well-established body of prior results, without evidence proportional to how extraordinary that claim actually is, gets held rather than run. And when a company's own blog post about a result uses language noticeably more confident than the paper underneath it, we treat that gap itself as the story, or as a reason to wait, rather than borrowing the blog's confidence for our own framing. We're also wary of pattern floods — the weeks where a dozen papers all apply the same trendy wrapper to a dozen unrelated problems, each claiming novelty for essentially the same idea. One or two of those are usually worth a mention. The other ten are noise wearing the same costume, and running all twelve as separate breakthroughs would flatter none of them and mislead a reader trying to gauge how much genuine progress actually happened that week.
Sometimes the Right Story Is Skepticism, Not Celebration
A meaningful share of what we end up publishing isn't "look at this new result," it's closer to "here's why a claim that's already traveled widely doesn't hold up under a closer read." We think debunking coverage deserves the same editorial weight as discovery coverage, maybe more, given how fast an unverified number can spread and start shaping a narrative, a funding decision, or a roadmap before anyone gets around to checking it. Being first to cover something impressive matters less to us than being right about whether it's actually impressive, and those two goals occasionally point in different directions. That tradeoff costs us something real — a competitor willing to run a claim uncritically will occasionally beat us to a headline — and we've made peace with that cost, because the alternative is optimizing for exactly the wrong thing on the exact stories where it matters most to get it right.
None of this is a formula, and we don't pretend our judgment is perfect on any given day, or that two editors here would always draw the line in exactly the same place on a borderline paper. But the alternative to a stated filter isn't neutrality. It's an unstated one, applied just as often and just as consequentially, with nobody able to check it against anything. We'd rather you know what ours actually is than assume everything carrying our name cleared some invisible bar we never wrote down. Read it, disagree with a call we made if you think we got one wrong, and hold us to the same standard we're applying to everyone else's abstract.