Peer review used to mean checking claims, baselines, and whether Figure 4 was secretly three charts in a trench coat. Now it also means asking a simpler question: does this citation exist on the same plane of reality as the paper? That is not glamorous work, but neither is unit testing, and somehow civilization continues. GeoSpatial ML’s Caleb Robinson and Isaac Corley put a concrete number on the problem in a July 30, 2026 post. According to GeoSpatial ML, they reviewed 22 machine learning conference submissions this summer and found 15, or 68 percent, with entirely fabricated citations, fabricated author lists for existing papers, or clear LLM generated writing. The useful lesson is not panic. It is that bibliographies have become a practical audit surface. ## What GeoSpatial ML found According to GeoSpatial ML, Robinson and Corley’s review assignments were spread across NeurIPS, WACV, and TerraBytes, a geospatial workshop at ECCV. The flaws they describe were not mere style complaints about prose that smelled faintly of autocomplete. They included fabricated citations, hallucinated authors, hallucinated technical jargon, nonsensical writing, and irrelevant citations. The authors called the work being in the “slop trenches,” which is unfortunately vivid and also sounds like a conference workshop nobody should attend before coffee. The important move is what happened after the complaint. GeoSpatial ML says Robinson and Corley released the paper references audit they now run, including a bib-audit skill created by Isaac Corley. That turns a fuzzy reviewer suspicion into something closer to quality control: verify that cited works exist, that author lists match, and that the reference actually supports the surrounding claim. In ML terms, the bibliography is now an evaluation set, except the labels are not supposed to be imaginary. ## Why citation audits beat vibe detection The arXiv paper listed as Phantom References: Hallucinated Citations That Survive Peer Review at Top Tier Conferences frames the same failure mode at a broader level: hallucinated citations can appear scholarly enough to pass through review. That matters because fake references are unusually audit friendly compared with vague prose. You do not need to divine whether a sentence was LLM assisted from its aura. You can check whether the named paper, authors, and cited claim line up. Rachel So’s paper, The Economics of AI Slop, gives the incentive story behind the mess. So argues that the marginal cost of generating a research paper with LLMs has fallen from thousands of dollars in researcher time to a few dollars of compute, creating pressure toward quantity over quality. In that setting, reviewers become the rate limiter, the poor little API endpoint getting hammered by cheap submissions. Citation audits are not a full defense against weak work, but they raise the cost of passing off synthetic scholarship as verified scholarship. ## What reviewers and authors should do now GeoSpatial ML’s post is most useful as a workflow nudge for researchers, not a dunk tank. Before submission, authors should audit references the way they audit experiments: confirm the cited paper exists, confirm the author list belongs to that paper, and confirm the citation is relevant to the sentence using it. If an LLM helped draft a related work section, the reference list deserves extra scrutiny, because models are famously confident librarians who sometimes shelve Atlantis under computer vision. For reviewers, the audit can be triage rather than a heroic weekend excavation. Start with citations that support central claims, unusually broad related work statements, and references that look oddly formatted or contextually irrelevant. If those fail, the review can focus on reproducibility and contribution with better evidence. This is not anti AI. It is pro provenance, which is peer review’s least sexy virtue and possibly its most useful one. ## The bigger publishing pressure Data & Society’s Ranjit Singh described arXiv as facing an influx of AI slop where surface polish competes with substance, and the essay notes that arXiv tightened moderation on October 31. That context matters because conferences and repositories are dealing with the same asymmetry: generating plausible text is cheap, while checking scholarship is slow. When the surface gets smoother, reviewers need more mechanical checks, not sharper vibes. The next thing to watch is whether citation auditing becomes normal conference hygiene. GeoSpatial ML’s small sample is not a census of ML publishing, and it should not be treated like one. But 15 flawed submissions out of 22 review assignments is enough to justify a new reflex: trust the PDF less, check the bibliography sooner. The model may write fluently, but the references still have to show ID at the door. ## Sources - Q&A from the slop trenches, GeoSpatial ML

Sources