Proofig AI expands beyond image integrity: meet Octym, the complete manuscript review →

Detecting AI-Generated Images in Research: Why Text Detection Isn’t Enough

Abstract illustration of scientific integrity with microscopy patterns and molecular visuals in cool tones.

AI-generated text detectors and AI-generated image detectors solve fundamentally different problems. A manuscript can pass text-based AI screening and still contain fabricated figures. Tools that analyze linguistic patterns have no mechanism for evaluating pixel-level image data. Researchers who screen only for AI-written text may still have an important gap in their integrity workflow.

AI-Generated Scientific Images Are Already Reaching Publication

This is not a hypothetical risk. In early 2024, a paper published in Frontiers in Cell and Developmental Biology contained AI-generated rat anatomy images with nonsensical anatomical labels. The figures passed peer review and were published before the paper was retracted by the publisher. The case demonstrated that AI-generated scientific figures can be convincing enough to clear editorial and reviewer scrutiny.

The broader landscape of image integrity problems provides important context. Image integrity specialists have reported that 20 to 35% of submitted life-science manuscripts are flagged for image-related issues during editorial screening at the submission stage. Not every flagged case reflects intentional wrongdoing. Errors in figure preparation and data handling can produce similar patterns, which is why screening is meant to surface issues for human review rather than to assign blame. Image-related problems are already a well-recognized challenge at the submission stage, even before AI-generated figures enter the picture.

AI-generated figures add a new dimension to this existing problem. As Nature has reported, generative AI tools can now produce convincing fraudulent scientific figures that are difficult to distinguish from real microscopy, histology, and other modalities. The tools are widely accessible, and the threat is recognized as systemic, not isolated.

Why Can’t Text Detection Tools Catch AI-Generated Images?

AI text detectors, such as those screening for ChatGPT or Claude output, work by analyzing linguistic patterns: token probabilities, sentence structure, stylistic uniformity, and statistical regularities in word choice. These methods are effective for their intended purpose. But they operate entirely in the domain of language. They have no capacity to evaluate pixel data, texture patterns, or structural anomalies in an image file.

This is not a limitation of any single vendor. It is a structural mismatch between the problem and the tool. An AI text detector examining a manuscript PDF will process the written content and skip over the figures entirely, or at most extract any embedded text labels. It cannot assess whether a microscopy image was captured by a real instrument or generated by a diffusion model.

The distinction matters practically. Many integrity screening workflows at the institutional level currently focus on text similarity and AI-written text detection. These are valuable checks. But a manuscript that clears them may still contain AI-generated figures that no part of the screening pipeline has evaluated. Researchers who assume their manuscript has been fully screened after text-based checks may be unaware of this gap.

The STM community has recognized this issue. Current industry-wide integrity tools have focused heavily on text, including detection of AI-generated nonsense text and plagiarism. Image-specific AI detection remains underserved across most screening platforms, creating a blind spot that affects researchers and publishers alike.

Why General AI Image Detectors Fall Short for Scientific Figures

Some researchers may wonder whether consumer-facing AI image detectors, tools designed to identify AI-generated photographs, artwork, or social media content, could fill this gap. They cannot, for a specific reason: these tools are trained on photographic and artistic image types that bear little resemblance to scientific figures.

A detector trained on portraits, landscapes, and digital illustrations has learned to recognize artifacts characteristic of those domains. It has not been exposed to the visual patterns of confocal microscopy, Western blots, FACS plots, histology slides, or cell plate images. Scientific image types have their own textures, noise profiles, resolution characteristics, and structural conventions. A generative model producing a synthetic Western blot leaves different artifacts than one producing a synthetic photograph.

Effective detection of AI-generated scientific images requires domain-specific training on real research images across multiple modalities. This includes:

Image Type Why Domain-Specific Training Matters
Microscopy (confocal, electron, light) Unique noise patterns, resolution characteristics, and structural features differ from photographic images
Western blots and gels Band patterns, background gradients, and lane structures require specialized pattern recognition
Histology Tissue architecture, staining patterns, and cellular morphology have domain-specific visual signatures
Cell plates Colony distribution, well patterns, and growth characteristics are unlike any consumer image type
Medical scans Modality-specific artifacts (CT, MRI, PET) require training on clinical imaging data
Animal imaging In-vivo fluorescence, bioluminescence, and anatomical imaging have distinct visual properties

Proofig’s AI-generated image detection covers all of these scientific image types, identifying images created by the most widely used generative models. The platform is continuously updated as new models emerge, because the detection challenge evolves with each new generation of image synthesis tools.

What Does Image-Specific AI Detection Actually Analyze?

Without disclosing proprietary technical details, it is useful for researchers to understand the general approach. Image-specific AI detection for scientific figures analyzes pixel-level artifacts, texture inconsistencies, and structural patterns that are characteristic of generative models. These patterns differ by image type, which is why domain-specific training matters.

For example, a generative model producing a synthetic histology image may introduce subtle regularities in tissue texture that differ from the natural variability seen in real stained sections. A fabricated Western blot may show band edges or background noise distributions that are statistically inconsistent with real gel electrophoresis output. These signals are invisible to text-based tools and unreliable for general-purpose image detectors.

Critically, detection requires training on large volumes of real scientific images. Proofig’s platform has been validated on hundreds of thousands of real images across the major scientific image types. This training base is what allows the system to distinguish genuine from synthetic figures with the specificity that scientific publishing demands.

The Practical Risk of Relying Only on Text Screening

For a researcher preparing a manuscript for submission, the practical risk is concrete. A manuscript may clear text-based AI screening and text similarity checks, then be flagged during a journal’s image integrity review process. Many leading journals now screen submitted figures for integrity issues, including AI-generated content. The Science family of journals adopted Proofig specifically to detect altered and problematic images before publication.

If AI-generated figures are detected after submission, the consequences range from revision requests to desk rejection. If they are detected after publication, the consequences escalate to correction or retraction, with significant financial and reputational impacts for the researchers involved. Nature has reported on the intense personal and career stress that researchers face when retracting papers, even when the underlying cause is honest error.

It is worth emphasizing that AI-generated images can enter a manuscript through multiple routes. A collaborator may use an AI-assisted tool without realising the output is not acceptable for publication. A team member may generate a placeholder figure and forget to replace it. A student may not fully understand which AI-assisted workflows are acceptable. Screening before submission helps surface these suspected issues for human review before they become a journal’s problem.

What to Look for in an AI-Generated Image Detection Tool

Researchers evaluating tools for pre-submission image screening should consider these criteria:

  • Scientific image coverage. Does the tool specifically cover the image types used in your field (microscopy, Western blots, histology, medical scans, cell plates, animal imaging)?
  • Domain-specific training. Is the tool trained on real scientific images, not general photography or artwork?
  • Model currency. Does the tool update its detection capabilities as new generative AI models are released?
  • Human review framing. Does the tool flag suspected issues for your review, rather than issuing automatic verdicts? The goal is to support your judgment, not replace it.
  • Low false-alarm rate. A tool that flags too many legitimate images wastes your time and erodes trust. Look for tools that have been validated on large datasets of real scientific images and demonstrate low false-positive rates across microscopy and Western blots. Proofig is validated on hundreds of thousands of real scientific images across modalities, with distinct detection logic for each image type to keep false-alarm rates low.
  • Broader integrity coverage. A tool that also screens for image duplication, manipulation, and plagiarism addresses the full spectrum of image integrity risks, not just AI-generated content.

Closing the Gap Before Submission

If your current pre-submission workflow includes text similarity screening and AI-written text detection but does not include image-specific AI detection, you have a gap. Closing it is a protective step, not a response to suspicion. It is the same logic that drives researchers to check references, verify data, and proofread before submitting. Proofig’s image integrity platform provides AI-generated image detection alongside duplication, manipulation, and plagiarism screening, covering the image-specific risks that text tools were never designed to address.

Related Items