← Field Notes

NeurIPS 2026 Decisions Are Out. Pangram’s Own Report Says the Version That Screened the Position Papers Called Up to 14.9% of AI-Polished Human Reviews “AI.”

StrictCite Field Notes  ·  26 September 2026
English  ·  Deutsch  ·  日本語  ·  한국어
Every factual claim below links to a primary source: the track chairs’ own post, the NeurIPS call for papers, Pangram’s own technical report, or a peer-reviewed or preprint paper. Where we could not find a primary source, we say so rather than repeat the claim.
Abstract

NeurIPS 2026 released final decisions on 24 September[2]. In the Position Paper Track, 178 of 969 submissions had already been desk-rejected in June on the basis of Pangram AI-detection scores[1], and the decision date brought a new wave of claims about what happened in the appeals, who the detector penalised, and why. We checked twelve of the most repeated claims against primary sources. Some hold. Several are wrong or incomplete, including four statements in our own 8 September Field Note[11], which we have now corrected. Two we could not verify at all. We also read two primary documents that most commentary has skipped: Pangram’s own July 2026 technical report[4], which benchmarks v3.3.2 (the exact version NeurIPS used) against its successor, and an ICML 2026 paper by Saha et al.[5] that discusses the NeurIPS rejections by name. Neither supports the loudest version of the backlash. Both point to the same narrower problem: text that a human wrote and an AI then polished, which is exactly what the track’s policy allows.

1. What changed on 24 September, and what did not

The Position Paper Track’s call for papers lists 24 September as the date of final decisions[2]. That is the date the track’s surviving submissions learned whether they were accepted. It is not the deadline for the 123 conditionally flagged papers. Those authors had until 15 June to submit a version-history dossier or be desk-rejected[1].

What has not changed is the public record. As of 26 September, the most recent posts on the NeurIPS blog are about affinity events (25 September) and financial assistance (16 September)[3]. The chairs’ 2 June post has no addendum. Nothing public says how many of the 123 dossiers were accepted, how many authors withdrew, or how many position papers the track accepted in the end. Any number you see for those is not from NeurIPS, and we have not found one with a source.

Figure 1. How the chairs sorted 969 position papers
Figure 1. How the chairs sorted 969 position papers  ·  Download: PNG · SVG

2. Twelve claims, checked

These are the claims we saw repeated most often in write-ups and summaries around the decision date, including in our own earlier piece. Each verdict is based on the source linked in the same row.

Claim What the primary source says Verdict
24 Sept was the final desk-rejection date for the 123 conditional papers. The evidence deadline was 15 June[1]. 24 Sept is final decisions for the whole track[2]. Wrong
Appeals were rejected when large text blocks appeared in single commits, or when authors disclosed Grammarly or DeepL. We found no source. Dossiers are “reviewed by the PPT team”[1] and no outcomes are public. Separately, a controlled test found that Grammarly edits produced 0 of 100 “AI” verdicts from Pangram[5]. Not verified
79 papers were rejected for scoring over 0.8 and being solo-authored. The criterion is a score ≥0.8 plus either an author who “has submitted multiple solo-authored papers, with at least one above this threshold” or an author with “at least one other desk reject”[1]. Being a solo author does not trigger it on its own. Wrong (we said it too)
22 papers were rejected for scoring over 0.5 after the author denied any AI use. The criterion is a score ≥0.5 where the “author declared that they did not use AI, or did not declare AI use”[1]. It also catches authors who left the declaration blank. Incomplete (we said it too)
The 178 desk rejections had no appeal. “A standard desk-rejection that is not subject to appeal under standard circumstances”[1]. The 123 are labelled “Desk Reject with Appeal”. Holds, with caveat
At default 250–350-word windows, 42.7% of submissions scored 90–100% AI. Table 1 and Table 3 of the chairs’ post: 42.7% ≥90%, 28.2% at 100%, 70.5% ≥50%[1]. Holds
The chairs shrank the windows to 100 words to push the flag rate down. The numbers are right: 42.7% fell to 12.7% at ≥90%. The motive is not the one the chairs gave. They said smaller windows reduce “over-claiming AI use” and published the loss in recall that comes with it (Table 2)[1]. Numbers hold, framing doesn’t
Independent researchers ran the chairs’ papers through Pangram and got 24%, 36%, 45% and 69%. One person did: Sergey Berezin, in a LinkedIn post on 3 June. He named the four papers and wrote, “I make no claim whatsoever regarding the authorship process behind these papers”[6]. His post does not say his own paper was desk-rejected. Holds, one source (we overstated his status)
Pangram penalises non-native writers. Stanford found 61.22% false positives on TOEFL essays. The 61.22% figure comes from Liang et al. (2023), which tested seven other detectors[7], not Pangram. On the same TOEFL essays, Pangram’s own report gives v3.3.2 0 false positives out of 89[4]. That is vendor data and has not been independently replicated. Wrong as applied to Pangram
AI use rose more than tenfold in another NeurIPS track. In Evaluations & Datasets, papers scoring ≥90% went from 0.8% (2025) to 9.3% (2026)[1]. Holds
Substack added Pangram for paying subscribers and coined “Claudefishing”. Chris Best coined the term on 21 July. The scan runs on text over 100 words and “will show an analysis only to those who request it”[8]. Best’s post does not limit it to paying subscribers. Partly wrong
The community mobilised in huge Reddit threads throughout September. We could not find a specific thread to cite, so we are not repeating it. Not verified
Figure 2. Changing the window size moved the ≥ 90% share from 42.7% to 12.7%
Figure 2. Changing the window size moved the ≥ 90% share from 42.7% to 12.7%  ·  Download: PNG · SVG

3. What we got wrong on 8 September

Our earlier Field Note[11] had four errors. We have corrected that page today and marked it as updated. We list them here because this piece criticises others for the same mistakes:

  • We said the 79-paper tier was “score above 0.8 + solo-authored”. The actual criterion is in the table above.
  • We said the 22-paper tier required an explicit denial of AI use. It also covers authors who made no declaration.
  • We called Sergey Berezin “a desk-rejected author”. His post does not say that.
  • We attributed reviewer complaints about “Claude-speak” to his post. We cannot find that phrase in the source we cited, and we have removed it.

4. What Pangram’s own report says about the version NeurIPS used

The chairs screened the track with Pangram v3.3.2[1]. On 29 July, Pangram Labs published the technical report for v4 on arXiv[4], and almost every table in it compares v4 with 3.3.2. As far as we know, it is the only published benchmark of the model that screened the track. It is vendor data, measured on benchmarks the vendor chose or public ones it ran itself, not on position papers. With that caveat, these are its numbers:

Measure (Pangram’s own report) v3.3.2 v4 Where
False positives, 1M+ human-written English texts0.0539%0.0041%§5.3
False negatives, 520,000 AI texts1.99%0.34%§5.2
False positives, 24,586 English-learner texts2 (0.0081%)1 (0.0041%)Table 9
… of which Liang et al. TOEFL essays (n=89)00Table 9
AI-polished text called “AI” (n=11,363)0.18%0.01%Table 4
AI-polished human peer reviews called “AI”, easy / hard subset14.9% / 4.5%1.4% / 2.5%Table 28
Substantially AI-edited text correctly called “Mixed” (n=14,990)5.54%55.01%Table 5
Mean error in estimated AI fraction, interleaved documents28.6 pts5.9 ptsTable 7

We read three things from this, and the first cuts against the backlash we reported earlier.

Purely human text is not where the risk is. At v3.3.2’s reported false-positive rate of 0.0539%, you would expect about 0.5 false positives across 969 papers, if position papers behave like the vendor’s test set. Nobody has shown that they do. But if the number is even roughly right, false positives on entirely human-written papers cannot explain an 18.4% desk-rejection rate. The explanation has to involve AI of some kind: either AI drafting, or the AI polishing that the policy allows.

The ESL claim, as usually stated, does not match Pangram’s data. On the same TOEFL essays that produced the widely cited 61.22% figure for other detectors, v3.3.2 had zero false positives. The report does show the vendor treats these as the hard cases. Describing a training design it rejected, it says failures “concentrated … in specific registers: academic writing and ESL essays”[4]. We found no independent test of v3.3.2 on academic writing by non-native authors.

Polished human text is where v3.3.2 was weakest. On human-written peer reviews that an AI then polished, v3.3.2 called 14.9% of the easy subset fully “AI”. That is ten times v4’s rate on the same split. The row the report labels Pangram 3.0 shows 3.0%, so the version NeurIPS used did worse on this split than both its predecessor and its successor. The report also calls 3.3.2’s behaviour on mixed text “polarized”: it prefers to label AI-assisted texts Human or AI rather than Mixed. This is the category the track’s policy explicitly allows (“copy-editing or similar peripheral changes”[1]), The chairs did test it. On 100-word excerpts from 10 FAccT 2022 papers, none of the permissible edits they tried was flagged[1]. Ten samples is a small test next to the vendor’s own figure from a larger benchmark.

Figure 3. Pangram 3.3.2 called up to 14.9% of AI-polished human reviews “AI”
Figure 3. Pangram 3.3.2 called up to 14.9% of AI-polished human reviews “AI”  ·  Download: PNG · SVG
Figure 4. The “61% ESL false positives” figure was not measured on Pangram
Figure 4. The “61% ESL false positives” figure was not measured on Pangram  ·  Download: PNG · SVG

5. The ICML paper that discusses the rejections by name

Saha, Juneja, Chaudhuri, Sajeevan, Shah and Pruthi’s paper Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable appears at ICML 2026[5]. It is the most careful source we found, and its conclusion is more nuanced than either side of the argument. Its main finding is that all five detectors it tested “misclassify a non-trivial fraction of LLM-polished reviews as AI-generated”. Pangram called 3.1% of AI-polished human reviews “AI”, not even “Mixed”. In its appendix:

  • Grammarly: after accepting every Grammarly suggestion on 100 human reviews, Pangram returned 89 “Human”, 11 “Mixed” and 0 “AI”.
  • Machine translation: after round-tripping reviews through German, Chinese or Japanese with Google Translate and DeepL, Pangram called 1.1% to 4.5% of them fully “AI”, depending on language and translator order.
  • Style-instructed polishing (“use British English”, “use short sentences”): 2.42% called “AI”.

The translation result matters for people who draft in another language. The chairs’ own policy table classes “Translation / backtranslation” as borderline permissible[1].

On NeurIPS itself, the authors are explicit that their numbers do not prove the rejections were wrong:

“Our ∼3% false positive rate on AI-polished (H-AI) reviews is best understood as a per-window rate, not a per-paper one … If window-level false positive rates were independent, then it is improbable that a paper which is only polished using LLMs would have more than 50% of their windows flagged … But one key caveat is that this reasoning assumes independence across windows from the same paper. An author’s individual writing style may induce correlated errors across windows … Per-window false positive rates should therefore not be extrapolated to whole papers in either direction.”[5]

We ran that argument with v3.3.2’s worse rate. If each window of a polished-only paper were independently flagged at 14.9%, the chance that more than half of a 30-window paper gets flagged is about one in a million. Independence is the assumption doing all the work. It says an author’s style has no effect from one paragraph to the next, even though consistent style is what makes someone’s writing recognisable. Nobody has measured it. The criterion most exposed to it is the 22-paper tier, which sets its bar at 0.5 rather than 0.8 or 0.9.

6. Two things in the chairs’ post we could not reconcile

Which window setting decided the rejections. Table 3 of the chairs’ post says 12.7% of papers (about 123) scored ≥90% with medium windows, and 42.7% (about 414) did with default windows. Table 5, the decision table, desk-rejects 77 papers at ≥0.9 with no other conditions[1]. On its face, 77 matches neither count, and the post does not say which scores Table 5 used. There may be a simple explanation, such as a different score definition or later exclusions, but it is not in the post. Authors cannot check a threshold when they do not know which score it was applied to.

The reviewer side. The same policy required reviewers to “commit to not using AI tools to write their reviews”[1]. The post describes screening every submission and does not describe any screening of reviews. For a sense of how well a commitment holds without enforcement: in a randomised experiment at ICML 2026, 22.5% of reviewers assigned a no-LLM policy said in a survey that they had used an LLM anyway[9]. That is a different venue and says nothing directly about NeurIPS reviewers. It does show that the attestation the chairs judged insufficient for authors is being relied on for reviewers.

7. What is still not public

  • How many of the 123 conditional papers submitted a dossier, how many were cleared, and how many withdrew.
  • How many position papers the track accepted on 24 September.
  • Which window setting and score definition the Table 5 thresholds used.
  • Whether desk-rejected authors were shown their own scores, and how their text was split into windows.
  • Any calibration of v3.3.2 on formal academic prose by non-native authors, from Pangram or anyone else.

We will update this page if the chairs publish any of these. If you are an affected author and have primary documentation you are willing to share, such as the notification text or a dossier outcome, [email protected] reaches us. We will not publish anything we cannot verify.

8. If you are affected, or submitting next cycle

The rejection is described as a standard desk rejection. The chairs’ post uses the phrase “a standard desk-rejection”[1] and does not describe it as a misconduct finding. Once the paper is no longer under review at NeurIPS, the usual rules for submitting elsewhere apply. Check the new venue’s own dual-submission and AI-use rules.

Declare, specifically. The 22-paper tier covered authors who declared no AI use and authors who declared nothing. A blank declaration was treated like a denial.

Keep a version history. The dossier the chairs asked for was a pre-AI checkpoint, a post-AI checkpoint and the final paper, plus an analysis showing the AI edits added no new substance[1]. They wrote: “We expect that in future years this kind of audit trail will become a default.” Overleaf history, Git commits or Google Docs version history all produce this as you work. It is very hard to build afterwards.

Know what the evidence says about your tools. In the one controlled test we found, full Grammarly correction produced no “AI” verdicts from Pangram. Machine translation of a whole review produced them 1.1–4.5% of the time[5]. Both are small tests on review-length text, not papers.

Watch for alternatives to detectors. One proposal, greCAPTCHA (Payan, Gyevnár, Kasirzadeh & Shah, 17 September), tests whether the named authors understand their own paper instead of scoring the prose. Its prototype reached an AUC of 0.90 in predicting actual authorship, with 31 researchers[10]. It is a small prototype, and we know of no venue using it.

9. Where citation checking fits, and where it doesn’t

StrictCite does not detect AI prose, and nothing on this page is an argument for our product over Pangram. They answer different questions. Pangram estimates who wrote a sentence, and as sections 4 and 5 show, even its maker’s benchmarks put the uncertainty in the polished middle. A reference either resolves to a real record in Crossref, OpenAlex or DBLP or it doesn’t, and you can check the answer yourself. That is the other NeurIPS desk-rejection mechanism this cycle[12], and it is the one you can fully check before you submit. Check a bibliography free at strictcite.com.

References

  1. [1] Lu, A., Lazar, S., & Rügamer, D. (NeurIPS Position Paper Chairs), with Hua, S. & Metcalf, K. AI-Generated Papers in the NeurIPS 2026 Position Paper Track. NeurIPS Blog, 2 June 2026. Tables 1–5 and the appeals section. blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track ↑
  2. [2] NeurIPS 2026 Call for Position Papers: dates (final decisions 24 September) and AI-use policy. neurips.cc/Conferences/2026/CallForPositionPapers ↑
  3. [3] NeurIPS Blog, 2026 Conference category (post index, checked 26 September 2026). blog.neurips.cc/category/2026-conference ↑
  4. [4] Glickenhaus, B., Thai, K., Russell, J., Masrour, E., Han, Y., Spero, M., & Emi, B. (2026). Pangram 4 Technical Report. arXiv:2607.27183, 29 July 2026. §4.1, §5.2–5.3, Tables 4, 5, 7, 9, 28. arxiv.org/abs/2607.27183 ↑
  5. [5] Saha, R., Juneja, G., Chaudhuri, D., Sajeevan, N., Shah, N. B., & Pruthi, D. (2026). Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable. ICML 2026. arXiv:2603.20450v2. §1 (“AI detection on full papers and a note on NeurIPS PPT desk rejections”), Appendix B. arxiv.org/abs/2603.20450 ↑
  6. [6] Berezin, S. (2026). We Shouldn’t Desk-Reject Papers Based on Unvalidated AI Detection. LinkedIn, 3 June 2026. linkedin.com/pulse/we-shouldnt-desk-reject…-orc6e ↑
  7. [7] Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. doi.org/10.1016/j.patter.2023.100779 ↑
  8. [8] Best, C. (2026). Against Claudefishing. The Substack Post, 21 July 2026. post.substack.com/p/against-claudefishing ↑
  9. [9] Kim, S. S. Y., Deng, W. H., Vaughan, J. W., Su, B., Su, W., Agarwal, A., Li, S., Jaggi, M., Goldstein, D. G., Shah, N. B., & Dudík, M. (2026). Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026. arXiv:2609.19420. arxiv.org/abs/2609.19420 ↑
  10. [10] Payan, J., Gyevnár, B., Kasirzadeh, A., & Shah, N. B. (2026). greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI. arXiv:2609.20481. arxiv.org/abs/2609.20481 ↑
  11. [11] StrictCite Field Notes. NeurIPS Desk-Rejected 178 Position Papers for Being “AI-Generated.” The Detector Flagged the Track Chairs Too. 8 September 2026, corrected 26 September 2026. strictcite.com/blog/neurips-2026-position-paper-pangram-ai-detection ↑
  12. [12] StrictCite Field Notes. NeurIPS Started Desk-Rejecting Papers Over References That Don’t Exist. strictcite.com/blog/neurips-2026-hallucinated-citations-desk-reject ↑