I think a big reason for this is that PDF editing tools are just pretty terrible.
PDFs are meant to be a presentation document format, not an editable document format. I have some experience editing PDFs, and it's difficult to make significant modifications. Editing small sections of text or replacing a date is fine, but once you try to start editing entire paragraphs or changing the layout, it quickly becomes unmanageable. The kerning looks weird, the spacing between words looks off, and when you look into it you realize the the sentence you're trying to edit is actually split across a dozen text objects. For these it becomes almost impossible to edit without destroying the spacing, so you just give up.
And the approved way to edit these PDFs is via Adobe Acrobat Pro, which costs $20/month (ugh), which I'm not willing to pay for.
So I think this is why we just don't see that many PDF forgeries. The editing tools just aren't very good. Maybe it's different now we can just ask Claude to edit our PDFs for us.
Good to see you change it from "PDFs are Presentation Document Format" => "PDFs are meant to be a presentation document format". Next time, don't forget edit log..
I agree that you won't find forged PDF documents in the wild with any type of frequency, but they are forged in private settings. Forged or doctored bank statements are relatively more common. FinCEN identifies altered bank statements alongside W-2s, tax returns, gift letters, and employment information as examples of documents involved in mortgage fraud [1]. In March 2026, a mortgage loan officer was even charged by federal prosecutors with altering and creating bank statements to make borrowers appear qualified for a mortgage [2].
Not quite the scenario described (it wasn't an academic paper), but https://mjg59.dreamwidth.org/73317.html is my experience of dealing with a forged PDF. Part of the reason it's rare is that it's actually very hard to do it well!
The flip side to this is that when you do come across a forged PDF (and I have!), they can be really, really effective. Nobody expects them. There are usually tells in the metadata, but... there don't have to be.
Offtopic but the damage caused to an entire generation by Gwern's nicotine article is truly astounding. Obviously he's not responsible for the actions of others and it was written with good intentions but wow.
I have no sense of gwen's readership numbers, but I find it unlikely that one rather erudite blog has anything to do with global nicotine consumption. People have been smoking for centuries. The vaping industry had guerilla marketing campaigns to target college kids.
No rain drop thinks it is responsible for the flood, but I am hard-pressed to imagine a blog article greatly swayed hearts and minds.
Yeah, once you remove the hazards inherent in cigarette and/or cigar smoking from the discussion, the only solid objection I've heard to nictotine consumption is that it's probably more addictive than caffeine.
I understand why personal vaporizer use was demonized, but -goddamn- that's going to be a public health disaster only an order of magnitude or two smaller than when folks started demonizing nuclear power back in the mid 1900's.
Yes. Libgen/Sci-Hub have minimal metadata requirements. And they especially do not require ISBNs/DOIs because there are an incredible number of documents out there with neither one. It would be absurd to ban all pre-1970 books, and lots of journals to this day don't bother issuing DOIs.
Gwern, in case you are reading this, there's a typo in:
> (In the rare cases anyone tampers with PDFs, it quickly turns into a technical morass and he-said-she-said, and requires several orders of magnitude more effort to prove than to do; consider the hoops epxperts had to jump through in the Craig Wright cases, or even just a landlord editing a contract—where the contract was done through a digital document timestamping service!)
Ok this one actually baffles me -- anyone have the energy to explain what I'm missing? The repeated concern about forged scientific papers doesn't really track, but that's what academic institutions are for, at least in part.
If someone were to share a preprint PDF that I don't believe, my instinct would be to double-check it against the institution associated with its publication. If it was never published then it's basically just a screenshot of a blog post.
Ohhh okay the opening makes more sense now. The blog above is publishing papers in form of blog posts as a protest against scientific publishing. I love the energy, but ruminating on how weird it is that so few people take advantage of the thing that you're weird for doing in the first place feels... obtuse?
>PDF forgery is striking because it’d be so easy to do: find a useful research paper, edit it in any of the many PDF utilities, upload anywhere, wait for people to copy it (as they do), then take down yours; now you have an authoritative peer-reviewed research paper floating around the Internet with no links to you, in the perfect crime. (Or better yet: upload it directly to Libgen/Sci-Hub and let everyone else redistribute it.)
Seen as (unsigned) PDFs are basically images of documents, it's wild to claim that there are no forgeries, anyone doing a good ole fake document, fake diploma, fake wire transfer, fake articles of incorporation, etc... is using a pdf for faking. A nigerian prince scam sending an attachment mentioning 15.000.000$ locked up in inheritance would constitute a PDF forgery.
Gwern is talk about editing an existing pdf. Less so about creating one from scratch.
> There is plenty of incompetence, fraud, and malice online, often in PDFs… but only new PDFs. I can’t think of a single fraud accomplished by editing a real PDF & just uploading it for Google Scholar etc. or where I’ve been burned by even mislabeling.
PDFs are meant to be a presentation document format, not an editable document format. I have some experience editing PDFs, and it's difficult to make significant modifications. Editing small sections of text or replacing a date is fine, but once you try to start editing entire paragraphs or changing the layout, it quickly becomes unmanageable. The kerning looks weird, the spacing between words looks off, and when you look into it you realize the the sentence you're trying to edit is actually split across a dozen text objects. For these it becomes almost impossible to edit without destroying the spacing, so you just give up.
And the approved way to edit these PDFs is via Adobe Acrobat Pro, which costs $20/month (ugh), which I'm not willing to pay for.
So I think this is why we just don't see that many PDF forgeries. The editing tools just aren't very good. Maybe it's different now we can just ask Claude to edit our PDFs for us.
[1] https://www.fincen.gov/mortgage-loan-fraud
[2] https://www.justice.gov/usao-mdfl/pr/licensed-mortgage-loan-...
No rain drop thinks it is responsible for the flood, but I am hard-pressed to imagine a blog article greatly swayed hearts and minds.
This one?
It seems to be listing the pros and cons of nicotine and is well-sourced...
What "damage" do you see caused by this article?
I understand why personal vaporizer use was demonized, but -goddamn- that's going to be a public health disaster only an order of magnitude or two smaller than when folks started demonizing nuclear power back in the mid 1900's.
Can this actually be done without an ISBN or DOI number?
> (In the rare cases anyone tampers with PDFs, it quickly turns into a technical morass and he-said-she-said, and requires several orders of magnitude more effort to prove than to do; consider the hoops epxperts had to jump through in the Craig Wright cases, or even just a landlord editing a contract—where the contract was done through a digital document timestamping service!)
Specifically "epxperts"
If someone were to share a preprint PDF that I don't believe, my instinct would be to double-check it against the institution associated with its publication. If it was never published then it's basically just a screenshot of a blog post.
Ohhh okay the opening makes more sense now. The blog above is publishing papers in form of blog posts as a protest against scientific publishing. I love the energy, but ruminating on how weird it is that so few people take advantage of the thing that you're weird for doing in the first place feels... obtuse?
Seen as (unsigned) PDFs are basically images of documents, it's wild to claim that there are no forgeries, anyone doing a good ole fake document, fake diploma, fake wire transfer, fake articles of incorporation, etc... is using a pdf for faking. A nigerian prince scam sending an attachment mentioning 15.000.000$ locked up in inheritance would constitute a PDF forgery.
> There is plenty of incompetence, fraud, and malice online, often in PDFs… but only new PDFs. I can’t think of a single fraud accomplished by editing a real PDF & just uploading it for Google Scholar etc. or where I’ve been burned by even mislabeling.