A member posts a photo of a flooded high street after last night's storm. Twenty minutes later, half the thread is arguing about whether it is real, three people have called it AI, and at least one of them is wrong. Now it is your problem. Learning how to moderate AI images turns out to be less about spotting fakes than most moderators expect, and much more about having a rule that survives an angry poster at 11pm. This guide covers the rule, the check, and the argument that follows.

The short version
Good AI image moderation starts with a decision, not a detector: name the harm you are preventing (deception, flooding, harassment, or craft), then pick one policy shape (ban, disclosure, containment, or open) and write it so a stranger can apply it without asking you. Verification is the supporting act. Check the poster and the context first, then run the original file through a metadata and error level analysis check, and treat what you find as a signal rather than a verdict. Most of the workload is social rather than technical, so decide in advance how you will handle members accusing each other. And if your community includes people in the EU, transparency rules under the AI Act start applying on 2 August 2026, which will make labelled AI content far more common.

Decide What You Are Actually Protecting Before You Ban Anything

"No AI images" is not one rule. It is four rules wearing the same coat, and picking the wrong one is why some communities relitigate this every single month. Before you write a word of policy, name the harm you are trying to prevent. Everything else follows from that one sentence.

There are four harms worth separating. Deception, where a generated picture is passed off as a real photo of a real event. Flooding, where cheap output buries the posts people actually joined for. Harassment, where generated images target a specific member. And craft, where a community exists to celebrate skill, and generated work changes what the space is for. A local news group and a watercolour group both write "no AI" into their rules, and they mean completely different things by it.

Naming the harm also tells you where to spend your energy. If deception is the problem, you need verification habits and a labelling norm. If flooding is the problem, you need volume limits, and forensics are close to irrelevant. If harassment is the problem, you need a fast removal path and an escalation route, and the question of whether the image is convincing barely matters at all.

17% of the largest subreddits had rules governing AI by late 2024, up from 8.4% sixteen months earlier, according to Cornell Tech research covering more than 300,000 communities

Pick One of the Four Policies That Actually Hold Up

Commit to one policy shape and say which one it is, in public, in the rules. In my experience, the communities that struggle most are running an unofficial fifth policy: vibes, applied inconsistently by whichever moderator happens to be awake.

Full ban. Easiest to state, most expensive to police, because enforcement depends on catching everything. It fits art, photography, and craft communities where the harm sits at the centre of what the space is for. Disclosure required. Generated content is welcome if it is labelled in the title, the flair, or a tag. It fits mixed-purpose communities, and it scales better than a ban because members do most of the labelling for you.

Containment. Generated work is allowed, but only inside a weekly thread, a dedicated tag, or a sister community. This is the compromise that keeps a vocal minority happy without reshaping the main feed. Open with a truth rule. Anything goes, as long as nothing is presented as a real photograph of a real event or a real person. Meme and hobby spaces often land here, and it is the only shape that stays enforceable when your queue is enormous.

The policy you can enforce on your worst week beats the policy you like best on your best one.

Write the Rule So a Stranger Can Follow It at 2am

A rule works when a brand new member can apply it without asking you. That means three things have to be on the page: what counts, what happens, and where to appeal. Miss any one of them and the rule quietly becomes a queue.

Start with what counts, because "AI image" is blurrier than it was two years ago. Fully generated pictures are the easy case. The arguments come from edits: generative fill, sky replacement, background removal, object cleanup, AI upscaling, and the enhancement your members' phones apply by default without telling them. Pick a line and write it down. One that works: "Generated or substantially altered images must be labelled. Routine adjustments such as cropping, colour correction, and noise reduction do not need a label."

Then put the consequence and the appeal route in the same breath. Something like: "A first unlabelled post is removed with a note. Repeats get a seven day mute. If we get it wrong, message the mod team and we will restore it." Consequences that escalate slowly leave you room to be wrong, and you will be wrong sometimes.

Before you publish the rule
Test the wording on three real posts from last month, including one you were unsure about at the time. If two moderators read the same sentence and reach different conclusions, the sentence is not finished. Keep rewriting until any disagreement is about the evidence rather than about what the rule means.

Once the wording holds up under that test, the rest of the job is checking. That part is faster than most moderators expect.

How to Moderate AI Images in 60 Seconds Without Becoming a Forensics Expert

Start with the poster, not the pixels. Context settles most cases faster than any tool will: how old is the account, has it posted anything else, is the same image landing in five groups within an hour, does the caption make a claim you can check against a news source, and does a reverse image search turn up an older copy of the same picture. If the story falls apart, you never need the technical step.

When context is not enough, ask for the original file. A screenshot has already been re-encoded and stripped of almost everything useful, so "please send the file straight from your camera roll" is the highest value sentence in AI image moderation. With an original JPG in hand, two checks cover most of the ground: metadata, which can show camera make and model, capture time, the software field that often names an editor or a generator, and any content credentials attached at creation; and error level analysis, which highlights regions of a JPEG compressed differently from their surroundings and can expose pasted or edited areas.

Image metadata panel with camera model highlighted during a moderation check
Image metadata

If you want both checks in one place, Fake Image Detector runs error level analysis and metadata analysis on an uploaded JPG in a few seconds, which is usually enough for a first pass on something sitting in your queue. Keep the habit cheap. A check that takes a minute gets used on a busy night. A check that takes fifteen gets skipped on exactly the night you needed it.

What Image Checks Can and Cannot Prove

This is the part most moderation guides skip, and it is the part that keeps you out of trouble. These tools produce evidence, not verdicts. A moderator who forgets that will eventually remove a real photographer's work, state publicly that it was generated, and lose a good member over it.

Missing metadata means almost nothing. Platforms re-encode images on upload and routinely strip embedded data, including content credentials, as a side effect of compression rather than as a policy choice. Screenshots, messaging apps, and any download-and-reupload cycle do the same. A bare file is the normal condition of an image that has travelled.

A file with no metadata is not a confession. On most platforms it is just a file that has been uploaded.

Error level analysis has the opposite shape. It is informative about editing and much weaker as a test of origin. It does its best work on an original JPEG that claims to be a straight photograph, where a pasted region can carry a different compression signature from the rest of the frame. A fully generated image saved once often looks uniform under ELA, and re-saving washes the signal out. Read a clean result as "nothing found here," never as "this is real."

The same caution runs in the other direction. Content credentials and platform AI labels are useful when they are present and meaningless when they are absent. Detector confidence scores are probabilities, not findings. Three weak signals pointing the same way make a case worth acting on. One tool output is a hunch with a percentage attached.

Handle the Accusation Before You Handle the Image

The image is usually the smaller half of the problem. Public accusations do the real damage, and they land hardest on the people least equipped to defend themselves: hobbyists whose work happens to be clean and well lit, illustrators with a smooth digital style, anyone whose holiday photo looks a little too good.

Two rules fix most of it. First, accusations go to the mod team rather than the thread, and calling another member's work AI in public becomes a violation of its own. Second, the burden sits with you, not with the accused. If you cannot support a claim, do not make one. "Removed pending source under rule 4" is defensible. "This is AI" is a statement of fact you may have to retract.

When you get it wrong, restore the post quickly and say so where the removal was visible. Being correctable in public costs a moment of pride and buys years of trust. Communities forgive mistakes far more easily than they forgive a mod team that never admits to one.

Some Posts Are Not a Rules Problem, They Are a Safety Problem

A small set of posts needs a separate track that is faster and quieter than everything else. Sexualised or degrading images of a member, impersonation of a real person, fabricated fundraising images after a disaster, and doctored evidence in a dispute between members all belong on it. None of these should ever sit in a queue waiting for a policy debate.

The workflow is short. Remove first and discuss afterwards. Preserve evidence before it disappears: the post link, the account name, timestamps, and a copy of the file. Report it to the platform rather than handling it alone, because account level action is beyond what any mod team can do. Then contact the person targeted, tell them what you removed, and point them toward takedown options instead of leaving them to work it out alone.

What Changes for European Communities in August 2026

If your group has members in the EU, this is worth ten minutes of your attention. The transparency rules in Article 50 of the EU AI Act start applying on 2 August 2026. They require providers of generative systems to mark synthetic output in a machine-readable way, and they require deployers to disclose deepfake content. Generative systems already on the market have until 2 December 2026 to meet the marking requirement, and content generated before August does not need to be labelled retroactively.

For a volunteer moderator, the effect is mostly indirect. The obligations sit with AI providers and with people deploying these systems professionally, not with an unpaid mod team running a hobby forum. What changes is the environment around you: labelled and machine-marked AI content becomes far more common, undisclosed generation starts to look deliberate rather than careless, and members will quote the rules at each other with wildly varying accuracy. Timelines differ elsewhere, so do not copy a European framing into a global group and present it as the law everywhere.

Quick check for groups with EU members
Add one line to your rules stating that platform AI labels and content credentials count as disclosure, and that stripping or obscuring a label is itself a violation. That keeps your policy pointed in the same direction as the labelling rules without turning your mod team into a compliance department.

Build the Workflow Once, Then Stop Rewriting It

Consistency beats accuracy over a long enough stretch. Members can live with a rule they disagree with. They cannot live with a rule that changes depending on who is on shift. So build the machinery once and let it carry the routine cases.

Four pieces cover most communities. An automated filter that flags posts containing generator names and phrases like "made with," so they reach the queue early. A required tag or flair for disclosed AI posts, which turns enforcement into something members can self-serve. Canned removal messages that quote the rule and give the appeal route, so nobody has to compose that message at midnight. And a shared decision log, even a single pinned document, recording borderline calls and what you decided, so your rulings stay consistent as the team turns over.

For example, if someone submits an AI-generated landscape and clearly marks it with the required "AI Generated" flair, the automated filter can verify the disclosure and send it through the normal review process. If another user uploads a similar image without disclosure, it is automatically queued for manual review, ensuring both cases are handled consistently under the same rules.

Then review the whole thing quarterly. Generator quality moves, platform labelling moves, and a policy written eighteen months ago is describing a problem your community no longer has.

FAQ

Can I just run every suspicious post through an AI detector?
Not on its own. Detectors return probability scores, and they are least reliable on exactly the cases you care about: heavily compressed uploads, screenshots, and images that mix real photography with generated elements. Use one as a single input alongside metadata, context, and the poster's history, and let several agreeing signals carry the decision rather than one number.
Should we ban AI images completely?
Only if the harm you identified is about craft or authenticity at the core of the community. Full bans are simple to write and expensive to enforce, because they depend on catching everything. Disclosure rules survive growth better in mixed communities, since members handle the labelling and you only deal with the exceptions.
What do I do about a member who was accused and turned out to be innocent?
Restore the post, correct the record where the accusation was visible, and make the correction as public as the removal was. Then look at whether your rules permit public accusations at all. Communities that keep having this problem are usually missing a rule against it.
Does the EU AI Act mean my group has to label AI images?
The Article 50 obligations fall on providers of generative AI systems and on deployers using them professionally, not on volunteer moderators of a community group. What changes for you is the surrounding environment: more labelled content, more machine-readable marking, and members who treat disclosure as the norm from August 2026 onward.

Start With One Rule and One Check

None of this asks you to become an image forensics specialist. It asks for one decision about what your community is protecting, one rule written plainly enough to apply at 2am, and one repeatable check you can run in under a minute.

Do the small version this week. Write the rule, pin it, and the next time something lands in your queue that feels off, ask for the original file and run it through Fake Image Detector before you decide anything. You will be right more often, you will be wrong more gracefully, and the arguments will get shorter every month.