← Field Notes

NeurIPS Desk-Rejected 178 Position Papers for Being “AI-Generated.” The Detector Flagged the Track Chairs Too.

StrictCite Field Notes  ·  8 September 2026
English  ·  Deutsch  ·  Русский  ·  日本語  ·  한국어
The Position Paper Track's review cycle is still ongoing. What we report below is what was publicly verifiable as of 8 September 2026, sourced from the track chairs' own blog post, the NeurIPS call for papers, and the community discussion that followed.
Abstract

In June 2026, the NeurIPS Position Paper Track screened all 969 submissions through Pangram, a proprietary AI-text-detection tool. 178 papers — 18.4% of the track — were desk-rejected outright, with no appeal process. Another 123 were told to provide version histories or face the same outcome. When independent researchers ran the track chairs' own recent papers through the same detector, scores came back between 24% and 69%. The tool used was version 3.3.2; Pangram 4, with its improved claimed accuracy, was not released until a month after the decisions were made. We traced every claim in this piece to a primary source — the track chairs' own published statement, the NeurIPS call for papers, and Pangram's own technical report — rather than repeating the community discussion at face value. This is a different mechanism from the hallucinated-citation checker we documented in our previous Field Note[7]: that one checks whether your references are real; this one checks whether your prose was written by a human. Both result in desk rejection. The false-positive risks are completely different.

1. What actually happened

The NeurIPS 2026 Position Paper Track, chaired by David Rügamer (LMU Munich), Alex X. Lu (Microsoft Research) and Seth Lazar (Johns Hopkins), required all submissions to be “substantially human-written,” permitting generative AI only for “copy-editing or similar peripheral changes”[1]. That is stricter than the main conference track, where the constraint is on hallucinated citations and undisclosed AI use, not on AI-assisted drafting itself.

To enforce it, the chairs partnered with Pangram, an AI-detection company, and ran every one of the track's 969 submissions through Pangram version 3.3.2[2]. The tool divided each paper into overlapping text windows and assigned an AI-probability score to each segment. Under Pangram's default windowing (approximately 250–350 words per segment), 42.7% of all submissions scored in the 90–100% AI-generated range, and 28.2% — 273 out of 969 papers — received a full 100% score[2].

The chairs then re-ran the detector using a smaller “medium” window of approximately 100 words, reasoning that narrower windows would reduce over-flagging of papers where AI use was minor or incidental. This dropped the 90–100% flagging rate from 42.7% down to 12.7%[2]. From the refined results, the chairs split submissions into three buckets:

Outcome Papers % of track
Desk-rejected, no appeal17818.4%
Conditional — evidence required by June 1512312.7%
Cleared66868.9%

The 178 desk rejections were not a single threshold. They were driven by three separate criteria[2]:

77 papers: Pangram score above 0.9
79 papers: Score above 0.8 + solo-authored
22 papers: Score above 0.5 + author explicitly denied any AI use

That last tier is worth reading twice: 22 papers were desk-rejected for scoring above fifty percent on an AI detector, because the author's own attestation said they used no AI at all, and the chairs treated the disagreement as evidence of dishonesty rather than a detection error. No appeal process was available for any of the three groups.

The 123 conditional papers were given until June 15, 2026 to submit version histories or document checkpoints showing the state of their paper before and after any AI-assisted edits. If they couldn't, or didn't, the paper was rejected[2]. The Position Paper Track's final notifications are not due until September 24, 2026[1], so the outcome of those 123 cases is not yet public.

2. How they tried to calibrate it, and what happened when others checked

The chairs did attempt validation. Before issuing rejections, they tested Pangram against a “negative control” set: papers from ACM FAccT 2022, a conference whose proceedings predate the widespread availability of ChatGPT[2]. On that pre-LLM corpus, the detector returned a 0.0% false-positive rate. That is a real, reasonable sanity check, and it is worth giving them credit for running it. The problem is what it doesn't test: whether the detector also returns 0.0% on 2026 position papers that happen to be written in the formal, structured, argumentative style that AI detectors are known to confuse with LLM output.

The community ran its own calibration test, and it was more damaging. Independent researchers took recent papers authored by the three track chairs themselves and ran them through the same Pangram detector. The reported AI-probability scores ranged from 24% to 69%[3].

This was not a gotcha for its own sake — it was a calibration test on the exact population the tool was applied to. If the detector cannot reliably clear the professional, peer-reviewed work of the people who chose it, the question is not whether those papers are AI-generated (they obviously are not), but what that tells you about the false-positive behavior on this population. Recall that 22 of the 178 rejections were issued for scores above 0.5. The chairs' own papers scored up to 0.69. Under their own criteria, at least one of their papers would have landed in the conditional-review bucket, and possibly the outright-reject pile, had it been a submission rather than a sanity check run by critics.

3. The circularity problem

Every NeurIPS submission includes an author attestation. For the Position Paper Track, authors attested their paper was substantially human-written. Here is how the enforcement then worked:

1. Author attests: “This paper is substantially human-written.”
2. Pangram scores the paper as AI-generated.
3. The attestation is deemed inconsistent with the score.
4. The inconsistency is treated as grounds for rejection.

The detector's output functioned as ground truth. If it disagreed with the attestation, the attestation lost. Critics called this circular: the tool's own score was the sole evidence used to contradict the only person who actually knows how the paper was written[3]. For the 178 outright rejections, there was no stage at which the author could present counter-evidence at all.

4. What Pangram's own numbers actually say

Pangram 4, released in July 2026, reports an AUROC of 0.9916 and a false-positive rate of 0.0041% — roughly 1 in 24,000[4]. These are vendor-reported figures from internal benchmarks. Two things make them inapplicable to the NeurIPS decisions:

First, NeurIPS used Pangram 3.3.2, not 4. The improved model was released a month after the desk rejections were already issued. Whatever Pangram 4's accuracy is, it is not the accuracy of the tool that made these decisions.

Second, even if the v4 figures applied, “1 in 24,000” is a rate measured on the vendor's own test distribution. Multiple independent studies have found that AI detectors' real-world performance on formal, structured academic prose — especially from non-native English speakers — diverges substantially from claimed accuracy[5]. The “simplicity” and formal rigidity of ESL writing statistically resembles LLM output, and a tool trained primarily on informal web text will see patterns that are not there. The track chairs did not publish a calibration study on the specific population of NeurIPS position paper authors, and we could not find one.

5. What the track chairs actually argued

The chairs' published rationale[2] is more interesting than the backlash suggests. Their central claim was not that AI-generated text is inherently bad, but that it “externalizes the cost of verification onto reviewers” — position papers are supposed to be original arguments, and if the argument was generated by an LLM, the reviewer is now doing the intellectual work the author was supposed to have done. They also raised credit attribution: if a model wrote it, who is the author?

They also published a number that makes the urgency real. Across the NeurIPS 2026 Evaluations and Datasets track, the number of submissions with a Pangram AI score ≥90% increased more than tenfold from 2025 to 2026[2]. That is not a Position Paper Track problem; that is a conference-wide trend, and the chairs were trying to get ahead of it. Reviewers in other tracks were independently describing a rise in what one called “Claude-speak” — prose that reads like a competent summary of an argument nobody actually made[3].

This is, whether or not you agree with how they enforced it, a real argument, and it is the same argument that runs through the hallucinated-citation problem we documented in our last Field Note[7]. In both cases, an automated system is used to detect something a human reviewer is unlikely to catch at scale, and in both cases the question is whether the detection is reliable enough to stake a desk rejection on.

6. How this compares to the citation checker

This is the second NeurIPS desk-rejection mechanism this cycle, and the two are easy to confuse. They should not be. They check different things, fail differently, and have completely different fairness profiles.

Citation checker (Field Note 05) Content detector (this piece)
What it checksWhether references exist in registriesWhether prose was written by a human
Claim typeVerifiable fact (“this DOI does not exist”)Probability estimate (“68% likely AI”)
False-positive sourcePDF noise, translations, pre-prints not yet indexedESL bias, formal academic style, calibration drift
Human review layers3+ (ACs → chairs → appeals) at ICLRNone for the 178; evidence window for the 123
Appeal processYes, at most venuesNo, for outright rejections
TrackMain track, four venuesPosition Paper Track only

The column that matters most is “Claim type.” A citation checker that says “this DOI does not resolve to any known work in Crossref, OpenAlex, or DBLP” is making a checkable factual claim. You can verify it. An AI detector that says “this paragraph is 73% likely to be AI-generated” is making a statistical claim about style. You cannot verify it by looking at a registry, because there is no registry. One is an audit; the other is an opinion with a confidence interval.

Neither is inherently wrong as a tool. But the standard for using either one to desk-reject a paper with no appeal should be, at minimum, that the tool's error rate on the actual population it is applied to has been measured and published. For hallucinated citations, the independent 48,000-paper audit we cited in our last piece[6] provides that. For Pangram on NeurIPS position papers, no equivalent study exists.

7. The part that should worry non-native English speakers most

In 2023, a Stanford study tested seven AI detectors on 91 TOEFL essays written by non-native English speakers. On average, 61.22% of those human-written essays were misclassified as AI-generated[8]. The same detectors were highly accurate on essays by native speakers. The mechanism is structural: non-native writers tend to use simpler vocabulary and more predictable syntax, which is statistically indistinguishable — to a model measuring “perplexity” and “burstiness” — from LLM output.

That study is three years old and tested earlier detectors. Pangram claims to have improved since. But the NeurIPS Position Paper Track used v3.3.2, and the chairs published no calibration data on the specific demographic breakdown of their submission pool. The international ML community submits heavily to NeurIPS. If even a fraction of the 178 rejected authors are non-native English speakers whose writing was misclassified, the fairness cost is not hypothetical.

One further detail: every submission was shared with Pangram under an enterprise-level data agreement guaranteeing that “zero data from the submissions would be retained” by the vendor[2]. Whether authors were informed in advance that their unpublished, anonymized research would be sent to a third-party commercial service is not clear from the published materials.

8. If you were rejected: what actually happens to you

This section exists because the people most affected by this are the ones most likely to be reading this article. Here is what we can verify about the consequences:

There is no blacklist. A desk rejection from the Position Paper Track is not a misconduct finding. It does not go on a permanent record. It does not trigger a NeurIPS Code of Conduct investigation. It does not ban you from future submissions. NeurIPS reserves the right to investigate and sanction actual scientific misconduct — plagiarism, fraud, prompt injection — but a desk rejection based on an AI-detector score is not classified as misconduct[1].

You can resubmit elsewhere. Once a paper is no longer under review at NeurIPS, you are free to submit it to another venue. You do not need to disclose the previous desk rejection. The relevant upcoming deadlines, based on current schedules:

Venue Deadline Notes
ICLR 202725 September 2026Imminent — check formatting requirements
ICML 2027~January–February 2027Official CFP not yet posted
AAAI 2027Passed (28 July 2026)Conference: 16–23 February 2027

There is no appeal channel for the 178. The desk rejections are final. The [email protected] address is specifically for Code of Conduct and harassment reports, not for paper disputes. Contacting the program chairs about a finalized desk rejection is unlikely to change the outcome.

If you were in the 123 conditional group: your deadline to submit version histories was June 15, 2026. Final track notifications are due September 24, 2026. If you provided the evidence, your paper should still be under review. If you did not, it was treated as a desk rejection.

Practical advice for next time: save every draft. Use tracked changes in your word processor or keep Git commits if you write in LaTeX. If you use any AI tool for anything — even Grammarly, even autocomplete — say so in your attestation. 22 of the 178 rejections were not for high AI scores; they were for the gap between a moderate score and a categorical denial. The attestation is not a quiz about what the detector will think. It is a disclosure requirement, and the safest answer is an honest, specific one.

9. What this actually means if you're submitting

This cycle, at NeurIPS alone, authors face two separate automated gatekeepers before a human reviewer ever reads their work. One checks whether your citations are real. The other checks whether your prose is “substantially human-written.” Both can end in a desk rejection. Neither is optional.

The citation-checking problem, at least, has a concrete defense: verify your own bibliography before you submit. Every claim is factual, every disagreement is checkable, and a tool that shows you the specific field-level conflict — this is what you wrote, this is what three registries say — gives you something to act on. That is the thing we built. The content-detection problem is harder, because there is no equivalent self-check: you cannot run your own paper through a proprietary detector whose windowing settings you don't know and whose calibration on your population has not been published. The best concrete advice anyone can give right now is: keep your drafts, save your version history, and if you use any AI tool for anything at all, say so in your attestation rather than claiming zero use — because 22 of the 178 rejections were for the gap between a detector's opinion and an author's denial, not for the AI use itself.

Check a bibliography for free at strictcite.com. For the other gatekeeper, write your own paper, keep your drafts, and hope the detector agrees.

References

  1. [1] Call for Position Papers. NeurIPS 2026. neurips.cc/Conferences/2026/CallForPositionPapers
  2. [2] Rügamer, D., Lu, A. X., & Lazar, S. NeurIPS 2026 Position Paper Track: AI-Use Policy and Screening Results. Published on neurips.cc. neurips.cc/Conferences/2026/PositionPaperTrackBlog
  3. [3] Community discussion, r/MachineLearning and X (formerly Twitter), June–July 2026. Multiple independent reports of track chairs' papers scoring 24–69% on the same Pangram detector.
  4. [4] Pangram 4 Technical Report. Pangram Labs, July 2026. Vendor-reported AUROC 0.9916, FPR 0.0041%. pangram.io
  5. [5] Who Wrote This? Evaluating the Reliability of AI Detection Tools in Higher Education. (2026). Comparative study of 160 documents across commercial detectors including Pangram, GPTZero, Copyleaks and Turnitin.
  6. [6] Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences. arXiv:2607.00738. arxiv.org/abs/2607.00738
  7. [7] StrictCite Field Notes. NeurIPS Started Desk-Rejecting Papers Over References That Don't Exist. strictcite.com/blog/neurips-2026-hallucinated-citations-desk-reject
  8. [8] Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. doi.org/10.1016/j.patter.2023.100779