1. The question nobody's marketing page answers
Ask any citation checker what it does and the answer is a version of the same sentence: does this DOI resolve, does this title match a registry record, is this paper real. That is a real, useful question, and it is not the one a fabricated co-author needs asked. An AI tool producing a citation rarely has to invent an entire fake paper — Attention Is All You Need is real, famous, and easy to get approximately right. What slips through unexamined is a name added to a real author list on a real paper with a real DOI. Every check built around “does this paper exist” passes it, because the paper was never the thing that was wrong.
We built our own free tool around exactly this gap[1] after testing whether the gap was real. This is that test, in full, including the parts where our own engine came up short.
2. What we actually tested
The base citation is the real, unedited BibTeX entry NeurIPS itself publishes for Attention Is All You Need[2]. From it we built six variants, each a single, isolated edit so a tool catching it is demonstrably catching that one thing:
- A fabricated ninth co-author added to the real eight.
- A fabricated DOI (confirmed non-resolving) attached to the real title and authors.
- An entirely invented paper, checked against Crossref, OpenAlex and a web search before use to confirm it collides with nothing real.
- The real DOI, with the year changed from 2017 to 2015.
- A real, famous retraction — Wakefield et al., 1998, The Lancet[3] — cited plainly, with no indication it was ever withdrawn.
- A chimera: the real Attention Is All You Need title, with BERT's real authors (Devlin, Chang, Lee, Toutanova) swapped in.
All six were run through CiteTrue, Citely, CiteMe, aicitationchecker.org and Sourcely — every tool tested chosen for having a genuinely free, no-signup-required entry point. Coverage was not perfectly even across all five, and we are reporting that honestly rather than smoothing it over: CiteTrue's own front end never surfaced a result for case 6 in this run; CiteMe's batch de-duplication logic collapsed four of the six “Attention is all you need”-titled variants into a single ambiguous comparison, leaving clean, individually-attributable results only for cases 3 and 5; and aicitationchecker.org's own input validator rejected case 3 outright before it could be checked at all — a finding in itself, covered in Section 4.
3. The six-case results
Read case 5 again: two of the three tools in this table matched the citation to a database record whose title literally starts with the word “RETRACTED,” and returned a plain, unqualified pass. Not a missing signal — a visible one, ignored. That is a more serious failure than case 1, because the evidence that something was wrong was already on screen.
CiteMe, run separately on the same six cases, correctly rejected the entirely invented paper (“No matching record found,” checked against six registries) and gave the retraction its own high-priority category — both genuine, fair points in its favor. But three of the six “Attention is all you need” variants collapsed into a single de-duplicated comparison in its batch report, and that merged comparison surfaced something worth reporting on its own: the “source record” it matched against carried DOI 10.65215/2q58a426 and a publication year of 2025 — the exact signature of the citation-spam scheme we documented in an earlier Field Note[4], five duplicate Crossref entries for this same paper under a fabricated 2025 date. CiteMe's own “corrected” suggestion for this citation was “Attention Is All You Need (2025)” — sourced from the fake entry, not the real one. Clicking “Use corrected” here would make a bibliography wrong, not right.
4. A tool that couldn't even read the citation
aicitationchecker.org's input validator rejected our invented-paper case outright: “Multiple publication years (3: 2019, 2019, 2019). If these are separate references, place each on its own line.” The citation is a single, normally-formatted reference to a venue whose own name happens to contain its year twice more — “Proceedings of the 2019 International Workshop on Efficient Deep Learning (WEDL 2019)” — a completely ordinary pattern for an annually-named conference or workshop. The validator appears to count year-shaped substrings in a line and assume three or more means multiple concatenated references, with no model of the fact that a venue's own name routinely repeats its year. That is worth reporting on its own terms: a citation checker that cannot ingest a normally-formatted reference to an annually-named venue before it ever gets to verifying anything.
5. Sourcely performed best. So we tried harder.
Sourcely caught five of six cleanly, with precise, specific reasoning — “Jacob Devlin et al. are the authors of BERT, not ‘Attention Is All You Need,’ which was authored by Ashish Vaswani and colleagues” is not a field diff, it reads like actual reasoning. That raised an obvious question: is it checking a live registry, or reciting one of the most-reproduced papers in any language model's training data? Attention Is All You Need has over 267,000 citations. Nothing about getting its author list right proves a tool is verifying rather than remembering.
So we repeated the sharpest three tests against a real paper published five weeks before this piece was written — Leak It: Per-Document Extraction Beyond Aggregate Membership Inference, arXiv:2608.00144, a single-author preprint by Victor Maricato, confirmed live[5] — on the reasoning that essentially nothing could have memorized it. Sourcely correctly identified the real sole author against both an added fake co-author and a full author swap, unprompted, from a paper with no meaningful presence anywhere a model's training data would have seen it. That is real evidence of live, registry-backed verification, not recitation, and it is the fairest, strongest result any tool produced in this whole test.
It also has a genuine, repeatable blind spot, and finding it took building the harder case rather than accepting the first result. A DOI that is real, resolves, and points to a completely different real paper — not a typo, not a placeholder, an active, working hijack — was called “Real” with no warning at all, twice, in isolation. Sourcely's own explanation text for both cases correctly said the DOI was wrong (“incorrectly points to…”); the verdict bucket stayed “Real” regardless. A follow-up round of ten more targeted edge cases mapped the actual shape of this: isolated bibliographic slips on an otherwise-identified record — a wrong page range, a fabricated issue number, reversed author order, a DOI that's obviously a one-character typo of the correct one — correctly and reasonably land in a middle “Uncertain” tier rather than a false “Fake.” That is sensible calibration, not a flaw. But an isolated DOI hijacked to a genuinely different real paper, with nothing else wrong, still got waved through as clean “Real” — the single most misleading thing a citation can do, because the link resolves and looks trustworthy precisely because it's real, just real for something else.
6. We ran the same twenty-two tests against ourselves
Anything less would make the rest of this piece a marketing document instead of a Field Note. We took the complete set — the original six, the three obscure-paper repeats, three more obscure-paper variants, and the ten-case follow-up battery — and ran all twenty-two through StrictCite's own engine, the same pipeline the workspace runs, with no offline shortcuts and every registry live.
The result we are proudest of: the exact DOI-hijack case that Sourcely, the strongest performer in this whole test, called “Real” twice. StrictCite followed the DOI, found what it actually resolves to, and returned a complete, sourced field comparison:
authors: manuscript “Maricato, V.” — datacite/openalex both say “Ruiz, Marco; Arana-Catania, Miguel; Ardila, David R.; Ventura, Rodrigo”
VERDICT: CONFLICT
Not “the DOI looks off” — the actual title and the actual authors sitting at that DOI, laid next to what was cited, disagreeing on both. The real, famous retraction got the clearest treatment of any tool tested, ours included in every other row of this comparison: verdict INTEGRITY, the most severe category the engine has, with every source's retracted title quoted directly. The fabricated co-author on both the famous and the obscure paper produced a clean field-level conflict naming all three independent sources that disagree with the manuscript, rather than a silent pass or a silent override.
It also found two things wrong with our own product, and we are naming both rather than only reporting on everyone else's:
- The BERT-authors chimera case, run with no DOI or arXiv ID present in the citation — title and author fuzzy-matching only — fell to UNVERIFIED instead of a clean CONFLICT: no candidate reached the match threshold, and the real paper didn't even appear among the near misses shown. The identical wrong-author scenario on the obscure paper, where an arXiv ID was present in the citation, resolved correctly and cleanly. A wholesale author swap should not be harder for the engine to explain than a single added name — if anything it's the more serious case — and right now it is, specifically when the citation gives us no identifier to anchor the lookup.
- On the venue-fabrication cases, the structured field data correctly shows the container/venue field as never parsed out of the citation at all — yet a separate message still fired claiming the fabricated journal name and “arXiv (Cornell University)” are “abbreviation variants of the same publication.” They are not, and the two halves of our own output disagreeing with each other is worse than either being wrong alone. This looks like a parsing gap: a venue name sitting between a title and a raw DOI/URL, in exactly the shape a fabricated-journal citation actually takes, isn't being recognized as the container field.
Neither ever changed a verdict from a safe one to an unsafe one — the engine never confidently confirmed a fabrication in either case, which is the property that actually matters — but both made the evidence shown less clear than it should be, and “less clear than it should be” is exactly the standard this whole piece is holding five other companies to. So we fixed them, and we are reporting exactly how far each fix actually reaches rather than rounding up.
The venue-parsing gap is closed. Root cause: a plain-text citation's container field had no rule for stopping at a trailing URL, so a fabricated journal name glued to the paper's own DOI link — “Nature Machine Intelligence. https://doi.org/10.48550/arXiv.2608.00144” — parsed as one blob, and the literal substring “arxiv” inside that URL tripped a filter meant to catch actual repository citations, discarding the fabricated venue as if it were absent. A URL now stops container extraction the same way a volume or year already did. Re-run on all three affected cases: each now correctly reports the venue as unconfirmed, with an accurate message instead of the false one.
The chimera gap is half closed, and we are naming the remaining half rather than letting the fix stand in for it. The actual bug — a guard built to stop a book review from supplying a real textbook's metadata was also blocking an exact title match, not just a similar one, from ever reaching the candidate the wrong author field needed to be compared against — is fixed and verified: a wholesale author swap with no identifier present now resolves to a clean, sourced CONFLICT, confirmed end to end on a real paper with a distinctive title. It does not yet fix the specific case in this piece. “Attention is all you need,” it turns out, is also the exact title of other unrelated real works on OpenAlex and Zenodo, and the canonical Vaswani paper is not reliably among the five results our own title search fetches per source for a phrase that generic. That is a separate, deeper retrieval limit, not the scoring bug we found and fixed, and we are not patching it in response to one worst-case title before weighing the trade-off properly.
7. What this is actually an argument for
Not that AI-based verification always fails on easy cases — Sourcely's performance on an obscure five-week-old paper is real, and pretending otherwise would be exactly the overclaiming this whole publication exists to catch other people doing. The argument is narrower and, we think, sturdier: the same category of problem got two different verdicts from the same tool depending on how plausible the fabrication looked, and that is a failure mode a fixed, auditable rule cannot have. A DOI either resolves to the cited work or it does not. That is not a judgment call a model makes more or less confidently case by case — it is a fact a rule checks the same way, every time, and shows you the receipt for.
Every claim in this piece is checkable against the same receipt: rule IDs, field-level diffs, and the exact DOIs and dates involved, not prose asserting a verdict. Run your own bibliography the same way: strictcite.com, or check a single co-author for free at the tool this test was built to justify: the Fake Co-Author Checker.