1. What actually happened
The NeurIPS 2026 Position Paper Track, chaired by David Rügamer (LMU Munich), Alex X. Lu (Microsoft Research) and Seth Lazar (Johns Hopkins), required all submissions to be “substantially human-written,” permitting generative AI only for “copy-editing or similar peripheral changes”[1]. That is stricter than the main conference track, where the constraint is on hallucinated citations and undisclosed AI use, not on AI-assisted drafting itself.
To enforce it, the chairs partnered with Pangram, an AI-detection company, and ran every one of the track's 969 submissions through Pangram version 3.3.2[2]. The tool divided each paper into overlapping text windows and assigned an AI-probability score to each segment. Under Pangram's default windowing (approximately 250–350 words per segment), 42.7% of all submissions scored in the 90–100% AI-generated range, and 28.2% — 273 out of 969 papers — received a full 100% score[2].
The chairs then re-ran the detector using a smaller “medium” window of approximately 100 words, reasoning that narrower windows would reduce over-flagging of papers where AI use was minor or incidental. This dropped the 90–100% flagging rate from 42.7% down to 12.7%[2]. From the refined results, the chairs split submissions into three buckets:
| Outcome | Papers | % of track |
|---|---|---|
| Desk-rejected, no appeal | 178 | 18.4% |
| Conditional — evidence required by June 15 | 123 | 12.7% |
| Cleared | 668 | 68.9% |
The 178 desk rejections were not a single threshold. They were driven by three separate criteria[2]:
79 papers: Score above 0.8 + solo-authored
22 papers: Score above 0.5 + author explicitly denied any AI use
That last tier is worth reading twice: 22 papers were desk-rejected for scoring above fifty percent on an AI detector, because the author's own attestation said they used no AI at all, and the chairs treated the disagreement as evidence of dishonesty rather than a detection error. No appeal process was available for any of the three groups.
The 123 conditional papers were given until June 15, 2026 to submit version histories or document checkpoints showing the state of their paper before and after any AI-assisted edits. If they couldn't, or didn't, the paper was rejected[2]. The Position Paper Track's final notifications are not due until September 24, 2026[1], so the outcome of those 123 cases is not yet public.
2. How they tried to calibrate it, and what happened when others checked
The chairs did attempt validation. Before issuing rejections, they tested Pangram against a “negative control” set: papers from ACM FAccT 2022, a conference whose proceedings predate the widespread availability of ChatGPT[2]. On that pre-LLM corpus, the detector returned a 0.0% false-positive rate. That is a real, reasonable sanity check, and it is worth giving them credit for running it. The problem is what it doesn't test: whether the detector also returns 0.0% on 2026 position papers that happen to be written in the formal, structured, argumentative style that AI detectors are known to confuse with LLM output.
The community ran its own calibration test, and it was more damaging. Independent researchers took recent papers authored by the three track chairs themselves and ran them through the same Pangram detector. The reported AI-probability scores ranged from 24% to 69%[3].
This was not a gotcha for its own sake — it was a calibration test on the exact population the tool was applied to. If the detector cannot reliably clear the professional, peer-reviewed work of the people who chose it, the question is not whether those papers are AI-generated (they obviously are not), but what that tells you about the false-positive behavior on this population. Recall that 22 of the 178 rejections were issued for scores above 0.5. The chairs' own papers scored up to 0.69. Under their own criteria, at least one of their papers would have landed in the conditional-review bucket, and possibly the outright-reject pile, had it been a submission rather than a sanity check run by critics.
3. The circularity problem
Every NeurIPS submission includes an author attestation. For the Position Paper Track, authors attested their paper was substantially human-written. Here is how the enforcement then worked:
2. Pangram scores the paper as AI-generated.
3. The attestation is deemed inconsistent with the score.
4. The inconsistency is treated as grounds for rejection.
The detector's output functioned as ground truth. If it disagreed with the attestation, the attestation lost. Critics called this circular: the tool's own score was the sole evidence used to contradict the only person who actually knows how the paper was written[3]. For the 178 outright rejections, there was no stage at which the author could present counter-evidence at all.
4. What Pangram's own numbers actually say
Pangram 4, released in July 2026, reports an AUROC of 0.9916 and a false-positive rate of 0.0041% — roughly 1 in 24,000[4]. These are vendor-reported figures from internal benchmarks. Two things make them inapplicable to the NeurIPS decisions:
First, NeurIPS used Pangram 3.3.2, not 4. The improved model was released a month after the desk rejections were already issued. Whatever Pangram 4's accuracy is, it is not the accuracy of the tool that made these decisions.
Second, even if the v4 figures applied, “1 in 24,000” is a rate measured on the vendor's own test distribution. Multiple independent studies have found that AI detectors' real-world performance on formal, structured academic prose — especially from non-native English speakers — diverges substantially from claimed accuracy[5]. The “simplicity” and formal rigidity of ESL writing statistically resembles LLM output, and a tool trained primarily on informal web text will see patterns that are not there. The track chairs did not publish a calibration study on the specific population of NeurIPS position paper authors, and we could not find one.
5. What the track chairs actually argued
The chairs' published rationale[2] is more interesting than the backlash suggests. Their central claim was not that AI-generated text is inherently bad, but that it “externalizes the cost of verification onto reviewers” — position papers are supposed to be original arguments, and if the argument was generated by an LLM, the reviewer is now doing the intellectual work the author was supposed to have done. They also raised credit attribution: if a model wrote it, who is the author?
They also published a number that makes the urgency real. Across the NeurIPS 2026 Evaluations and Datasets track, the number of submissions with a Pangram AI score ≥90% increased more than tenfold from 2025 to 2026[2]. That is not a Position Paper Track problem; that is a conference-wide trend, and the chairs were trying to get ahead of it. Reviewers in other tracks were independently describing a rise in what one called “Claude-speak” — prose that reads like a competent summary of an argument nobody actually made[3].
This is, whether or not you agree with how they enforced it, a real argument, and it is the same argument that runs through the hallucinated-citation problem we documented in our last Field Note[7]. In both cases, an automated system is used to detect something a human reviewer is unlikely to catch at scale, and in both cases the question is whether the detection is reliable enough to stake a desk rejection on.
6. How this compares to the citation checker
This is the second NeurIPS desk-rejection mechanism this cycle, and the two are easy to confuse. They should not be. They check different things, fail differently, and have completely different fairness profiles.
| Citation checker (Field Note 05) | Content detector (this piece) | |
|---|---|---|
| What it checks | Whether references exist in registries | Whether prose was written by a human |
| Claim type | Verifiable fact (“this DOI does not exist”) | Probability estimate (“68% likely AI”) |
| False-positive source | PDF noise, translations, pre-prints not yet indexed | ESL bias, formal academic style, calibration drift |
| Human review layers | 3+ (ACs → chairs → appeals) at ICLR | None for the 178; evidence window for the 123 |
| Appeal process | Yes, at most venues | No, for outright rejections |
| Track | Main track, four venues | Position Paper Track only |
The column that matters most is “Claim type.” A citation checker that says “this DOI does not resolve to any known work in Crossref, OpenAlex, or DBLP” is making a checkable factual claim. You can verify it. An AI detector that says “this paragraph is 73% likely to be AI-generated” is making a statistical claim about style. You cannot verify it by looking at a registry, because there is no registry. One is an audit; the other is an opinion with a confidence interval.
Neither is inherently wrong as a tool. But the standard for using either one to desk-reject a paper with no appeal should be, at minimum, that the tool's error rate on the actual population it is applied to has been measured and published. For hallucinated citations, the independent 48,000-paper audit we cited in our last piece[6] provides that. For Pangram on NeurIPS position papers, no equivalent study exists.
7. The part that should worry non-native English speakers most
In 2023, a Stanford study tested seven AI detectors on 91 TOEFL essays written by non-native English speakers. On average, 61.22% of those human-written essays were misclassified as AI-generated[8]. The same detectors were highly accurate on essays by native speakers. The mechanism is structural: non-native writers tend to use simpler vocabulary and more predictable syntax, which is statistically indistinguishable — to a model measuring “perplexity” and “burstiness” — from LLM output.
That study is three years old and tested earlier detectors. Pangram claims to have improved since. But the NeurIPS Position Paper Track used v3.3.2, and the chairs published no calibration data on the specific demographic breakdown of their submission pool. The international ML community submits heavily to NeurIPS. If even a fraction of the 178 rejected authors are non-native English speakers whose writing was misclassified, the fairness cost is not hypothetical.
One further detail: every submission was shared with Pangram under an enterprise-level data agreement guaranteeing that “zero data from the submissions would be retained” by the vendor[2]. Whether authors were informed in advance that their unpublished, anonymized research would be sent to a third-party commercial service is not clear from the published materials.
8. If you were rejected: what actually happens to you
This section exists because the people most affected by this are the ones most likely to be reading this article. Here is what we can verify about the consequences:
There is no blacklist. A desk rejection from the Position Paper Track is not a misconduct finding. It does not go on a permanent record. It does not trigger a NeurIPS Code of Conduct investigation. It does not ban you from future submissions. NeurIPS reserves the right to investigate and sanction actual scientific misconduct — plagiarism, fraud, prompt injection — but a desk rejection based on an AI-detector score is not classified as misconduct[1].
You can resubmit elsewhere. Once a paper is no longer under review at NeurIPS, you are free to submit it to another venue. You do not need to disclose the previous desk rejection. The relevant upcoming deadlines, based on current schedules:
| Venue | Deadline | Notes |
|---|---|---|
| ICLR 2027 | 25 September 2026 | Imminent — check formatting requirements |
| ICML 2027 | ~January–February 2027 | Official CFP not yet posted |
| AAAI 2027 | Passed (28 July 2026) | Conference: 16–23 February 2027 |
There is no appeal channel for the 178. The desk rejections are final. The [email protected] address is specifically for Code of Conduct and harassment reports, not for paper disputes. Contacting the program chairs about a finalized desk rejection is unlikely to change the outcome.
If you were in the 123 conditional group: your deadline to submit version histories was June 15, 2026. Final track notifications are due September 24, 2026. If you provided the evidence, your paper should still be under review. If you did not, it was treated as a desk rejection.
Practical advice for next time: save every draft. Use tracked changes in your word processor or keep Git commits if you write in LaTeX. If you use any AI tool for anything — even Grammarly, even autocomplete — say so in your attestation. 22 of the 178 rejections were not for high AI scores; they were for the gap between a moderate score and a categorical denial. The attestation is not a quiz about what the detector will think. It is a disclosure requirement, and the safest answer is an honest, specific one.
9. What this actually means if you're submitting
This cycle, at NeurIPS alone, authors face two separate automated gatekeepers before a human reviewer ever reads their work. One checks whether your citations are real. The other checks whether your prose is “substantially human-written.” Both can end in a desk rejection. Neither is optional.
The citation-checking problem, at least, has a concrete defense: verify your own bibliography before you submit. Every claim is factual, every disagreement is checkable, and a tool that shows you the specific field-level conflict — this is what you wrote, this is what three registries say — gives you something to act on. That is the thing we built. The content-detection problem is harder, because there is no equivalent self-check: you cannot run your own paper through a proprietary detector whose windowing settings you don't know and whose calibration on your population has not been published. The best concrete advice anyone can give right now is: keep your drafts, save your version history, and if you use any AI tool for anything at all, say so in your attestation rather than claiming zero use — because 22 of the 178 rejections were for the gap between a detector's opinion and an author's denial, not for the AI use itself.
Check a bibliography for free at strictcite.com. For the other gatekeeper, write your own paper, keep your drafts, and hope the detector agrees.