A 2006 paper in Nature helped steer nearly two decades of Alzheimer's research, hundreds of millions in funding, and a wave of drug development. It was retracted in 2024 after investigators concluded its key protein images had been spliced, duplicated, and partly erased. Fake images in scientific papers are not a fringe problem, and the story of how they get caught is odder than most people expect. Peer reviewers rarely find them. Journals often do not either. Here is how the fakes get in, and who actually stops them.
Why Peer Review Almost Never Looks at the Pixels
Peer review was built to check reasoning, not pixels. A reviewer is an unpaid specialist reading a PDF in a stolen hour, asking whether the experiment answers the question and whether the statistics hold up. Almost nobody zooms to 400 percent on a western blot to compare the speckle in the background of two panels. That gap is exactly where image manipulation lives.
That study asked 831 medical students and 26 working researchers to find duplications hidden in western blots, cell cultures, and tissue sections. Out of 34 duplications, the researchers found a median of 11. These were people who knew they were being tested and were actively looking. Reviewers reading a real manuscript, with no warning that anything is wrong, do worse.
Peer review asks whether an experiment makes sense. It was never designed to ask whether the picture of that experiment is real.
The Four Routes a Fake Figure Takes Into a Journal
Nearly every problem figure arrives by one of four routes, and only the last two are clearly fraud. Sorting them matters, because journals treat them very differently.
Beautification. A researcher erases a smudge, crops out a lane that did not work, or pushes contrast until a faint band looks convincing. Most people doing this think of it as tidying up. The rules say otherwise, because adjustments have to apply to the whole image, not to the part you wish looked better.
Reuse. The same control panel appears twice under two different labels, or a blot from an older paper reappears in a new one. Sometimes this is a copy-paste error while assembling a figure at 2am. Sometimes the control experiment was never run.
Splicing and cloning. Lanes from separate gels are stitched into one image without marking the join, or a patch of cells is cloned to fill an empty spot. This takes deliberate effort in an image editor.
Fabrication. The figure depicts an experiment that never happened. Paper mills sell these by the batch, often built from templates, and generative AI has made producing a plausible-looking blot a matter of minutes.
The proportions are known, at least for duplication. In Elisabeth Bik's landmark screen, of the 782 problem papers she identified, about 29 percent contained simple duplications, 46 percent contained duplications that had been repositioned, and 25 percent contained duplications that had been altered as well. That last quarter is hard to explain as a filing mistake. Moving and editing a panel takes intent.
How Common Are Fake Images in Scientific Papers?
About one in 25, based on the largest published visual screen anyone has done. Bik and her colleagues Arturo Casadevall and Ferric Fang inspected images in 20,621 biomedical papers across 40 journals published between 1995 and 2014. They found problems in 782 of them, or 3.8 percent, and judged that at least half showed signs of deliberate manipulation rather than error.
Rates climb when you point better tools at a narrower target. Sleuth Sholto David examined every paper with relevant images in the journal Toxicology Reports and found inappropriate duplications in 115 of 715 papers, around 16 percent. Software found things he had missed: 41 of those 115 slipped past his manual review and only turned up when he ran the papers through ImageTwin.
Pre-publication numbers look different, and that is the point of screening. When the American Society for Microbiology ran ImageTwin across 2,627 accepted manuscripts over a year, roughly 3.9 percent had duplications caught before publication. Most were sorted out with the authors. Six papers, about 0.23 percent, lost their acceptance outright. Sloppiness is common. Fraud is a minority of cases, but a minority of a very large number is still a lot of papers.
The Volunteer Sleuths Who Find Most of It First
Here is the uncomfortable part: once a paper is published, the people most likely to catch a doctored figure are not employed to do it. They are scientists working evenings, posting concerns on PubPeer, a public site where anyone can raise documented issues with a published paper.
Sholto David, a biologist in Wales, had flagged issues on more than 2,000 papers before most people had heard his name. On 2 January 2024 he published a post detailing suspect images across dozens of papers from senior researchers at the Dana-Farber Cancer Institute. Within three weeks the institute said it was requesting six retractions and 31 corrections. In December 2025 Dana-Farber settled a False Claims Act case for 15 million dollars, acknowledging that images and data had been misrepresented or duplicated in support of grant applications. David, as the whistleblower, received 2.63 million dollars.
The Alzheimer's case followed the same shape. Neuroscientist Matthew Schrag noticed anomalies, image analysts including Bik and Jana Christopher agreed, a Science investigation published the findings in 2022, and the retraction landed in 2024. In my experience reading these cases, the pattern almost never starts with a journal's own systems. It starts with one stubborn person and a zoom function.
How Journals Screen Figures Before They Publish
Journals have been screening images for longer than most readers assume, but only recently at scale. The Journal of Cell Biology began checking every accepted paper's figures back in 2002. Its long-running numbers were sobering: around a quarter of accepted manuscripts contained at least one figure that broke the rules and had to be remade, and about 1 percent had acceptance revoked because the manipulation changed what the data appeared to show.
Modern screening is software-assisted. Tools like Proofig and ImageTwin carve a figure into its individual panels, then compare every panel against every other panel in the manuscript, and against a database of previously published images. One vendor says its comparison set holds more than 150 million images from open-access biomedical literature. Crucially, the matching survives cropping, rotation, mirroring, and contrast changes, which is precisely where human eyes fail.
Adoption moved fast after 2023. The Science family of journals rolled Proofig out across all six of its titles in 2024. ASM adopted ImageTwin across its journals. MDPI signed a multi-year agreement with Proofig in 2025. Publishers also pool intelligence through the STM Integrity Hub, which by late 2025 reported around 40 publishers running more than 125,000 submissions a month through shared screening and intercepting roughly 1,000 suspected paper mill papers monthly. A report commissioned by STM in January 2026 described some publishers now staffing research integrity teams of more than 100 people.
What the Screening Software Still Misses
Duplication detection finds copies. It cannot tell you that a unique image is a lie. If a fabricated blot is generated once and never reused, there is nothing for the software to match it against, and the manuscript sails through the automated stage looking spotless.
Generated figures are the harder problem. A 2025 study in PeerJ tested three free web-based AI detectors on 48 western blots created with ChatGPT and DALL-E 3, alongside 48 authentic blots from papers published before generative AI existed. The detectors performed poorly, well below what you would want before accusing anyone of anything. Specialist vendors report much better results on microscopy and cell images, though independent verification of those claims is still thin.
The clearest illustration is also the most famous. In February 2024 a journal published a review article containing AI-generated illustrations of rat anatomy with impossible proportions and labels that were not real words. It was retracted three days later. Readers on social media caught it, not the two reviewers and not the editor. If figures that obvious can pass review, realistic ones already have.
The Forensic Checks You Can Run on a Figure Yourself
You can get further than you would think with a browser, a zoom function, and two free forensic techniques, as long as you accept one limitation up front: a published figure has been through a compression and layout pipeline that destroys some of the evidence. What you produce is a lead, not a verdict.
Work through it in this order. Download the highest resolution version available, usually the figure file from the publisher rather than a screenshot of the PDF. Zoom to 300 or 400 percent and compare the background of panels that are supposed to show different samples, because identical noise in two "different" images is the single strongest signal there is. Look along panel edges for rectangular seams, sudden brightness steps, or a hairline where two gels were joined. Scan for texture that repeats, which is what cloning leaves behind. Then reverse image search the panel to see whether it appeared in an earlier paper under a different label.
Error level analysis and metadata inspection are the two techniques worth adding, and both take seconds. ELA highlights regions of a JPG that were saved at a different compression history from their surroundings, which is how a pasted-in band sometimes reveals itself. Metadata can show editing software, timestamps, and export history, which occasionally survives in supplementary files even when it has been stripped from the main figure. You can run both on any JPG at Fake Image Detector: upload the figure, read the ELA map and the metadata report, and treat anything you find as a starting point for a proper question to the authors.
A caveat I would rather state than hide: publisher pipelines re-encode figures repeatedly, so ELA on a journal-hosted image is noisier than ELA on an original camera JPG, and a bright region is not proof of anything on its own.
As a practical example, imagine a western blot where one lane appears unusually bright in the ELA output. That could indicate a pasted band, but it could just as easily be the result of the journal compressing one section differently during production. Without the original image, neither explanation can be ruled out.
What Happens After a Paper Gets Flagged
Flagging a paper starts a slow ladder, not a switch. The journal asks the authors for the original, unprocessed data behind the figure. If the raw files exist and the error was honest, the outcome is a correction. If the authors cannot produce the originals, or the explanation does not hold, the journal may publish an expression of concern while it investigates. Retraction follows when the manipulation affects the conclusions. In parallel, the authors' institution runs its own inquiry, and funders such as the US Office of Research Integrity or the UK Research Integrity Office can get involved.
The timescales are brutal. The Alzheimer's paper was published in 2006, publicly questioned in 2022, and retracted in June 2024, by which point it had been cited roughly 2,300 times. Its retraction notice pointed to splicing, duplication, and use of an eraser tool. Retractions overall have surged, with more than 10,000 papers pulled in 2023 alone, but a paper keeps being cited the entire time a case grinds along, and citations rarely stop after retraction either.
Why a Doctored Blot Can Reach Your Medicine Cabinet
Fabricated figures are not a closed academic problem. They redirect funding, because grant committees back what looks promising. They shape which drug targets get chased and which clinical trials recruit patients. They also leak outward, into health journalism, patient advice, and supplement marketing that cites a paper long after it has been retracted.
It is worth being precise rather than dramatic here. The amyloid hypothesis in Alzheimer's did not rest on that single 2006 paper, and researchers still disagree about how much damage its retraction did to the wider field. What is not in dispute is that a specific protein assembly attracted years of follow-up work on the strength of images that had been altered. That is real money, real careers, and real time that patients did not have.
Every figure in a paper is data, not decoration. The moment it gets treated as illustration, fraud gets a free pass.
If You Spot Something Suspicious, Here Is the Move
Document first, accuse never. Take screenshots at high zoom with the figure and panel numbers visible, draw boxes around the matching regions, and write a plain description of what you observe: two panels labelled as different conditions share an identical region of background. Then raise it on PubPeer or email the journal's editorial office with the images attached. If nothing moves, the research integrity officer at the authors' institution is the next step.
A common mistake I notice is treating a forensic tool's output as the argument. A colourful ELA map convinces nobody at a journal, because editors have seen compression artefacts misread a hundred times. The claim that actually moves an investigation is a visual match a reader can verify in five seconds, followed by a request for the original data. Lead with that, and keep the pixel forensics as supporting material.
Questions Readers Ask About Image Fraud in Research
Are most image problems in papers actual fraud?
Can I tell whether a figure was generated by AI?
Does error level analysis prove a figure was manipulated?
What should I do if a paper I cited gets retracted?
Where This Leaves You
Fake figures get into journals because the system that reviews science was built to check ideas, not images, and it is only now catching up with software that compares panels at a scale no human can. The people who find the rest are volunteers with good monitors and unusual patience. Neither line of defence works quickly, which means a share of what you read today will be corrected years from now.
The practical takeaway is smaller than the problem. Next time a figure in a paper you rely on looks a little too clean, spend two minutes on it: zoom to 400 percent, compare the backgrounds of panels that should be different, then run the figure through Fake Image Detector for an ELA map and a metadata read. Most of the time you will find nothing, and you will read the next paper with sharper eyes anyway.