What AI image generators still get wrong about science
AI image models can produce a convincing photograph of a laboratory and cannot produce a correct benzene ring, and the difference is not about prompt quality — it is that one subject is visual and the other is factual.
Published 19 August 2026
The split that predicts everything
Every disappointment with an AI-generated science image comes from the same confusion, so it is worth stating plainly: image models render appearance, not fact. They have learned what a laboratory looks like. They have not learned that carbon forms four bonds.
That single distinction predicts the result before you type anything. A photograph of glassware on a bench cannot state a false fact — there is no fact in it. A structural formula is nothing but facts, and every one of them is a chance to be wrong.
So the reliable test is: could this image be incorrect, as opposed to merely ugly? If yes, generating it is the wrong tool. If no, generate away.
What they are genuinely good at
Laboratory scenes, glassware, crystals, colour changes, dissolution, flames, long-exposure light trails, telescope-style space imagery, abstract molecular decoration. Anything photographic, and anything abstract enough that there is no correct answer to get wrong.
These are not consolation prizes. A well-prompted lab bench beats most stock photography, and it costs nothing per image, which matters when you need forty header images for forty articles.
What they cannot do, and will not soon
Molecular structures. A diffusion model has no representation of valence. It will give carbon five bonds, close a six-membered ring with seven vertices, and label an atom with a misspelling. The image looks like a structural formula, which is precisely what makes it dangerous.
Anything with text. Element symbols, subscripts, axis labels, units. Text rendering has improved a great deal for short English words and is still unreliable for exactly the short technical strings science needs.
Charts and data. The model has no access to your numbers, so any chart it draws is invented — bars that do not match their own scale, axes that do not correspond to anything.
Diagrams where geometry means something. Free-body diagrams, vector arrows, ray diagrams. Arrow length and direction carry the meaning, and the model places arrows decoratively.
The workflow that actually works
Generate the atmosphere and add the information. Ask for the style frame with no text, no labels in the prompt, then place the real chart, the real structure and the real labels on top in a design tool.
You get a consistent visual language across a whole series without a single invented number, and it is faster than arguing with the model about a label it will never spell correctly.
For structures specifically, use a chemical drawing tool — ChemDraw, MarvinSketch, or the free RDKit. These build the image from the formula, so they cannot be wrong about it in the way a generated image always can.
Why this matters more in teaching material
An adult reading a marketing page can shrug off a decorative molecule. A student revising from a labelled diagram cannot, because the entire reason they are looking at it is that they do not yet know what it should say.
A generated diagram in teaching material is a diagram nobody has checked, presented to the one audience least equipped to catch the error. If a student might memorise it, it must not have been generated.
Common questions
- Can AI image generators draw accurate chemical structures?
- No. Image models have no representation of bonding rules, so they produce structures that look plausible and are chemically impossible — incorrect bond counts, wrong ring sizes, misspelled element symbols. Use a chemical drawing tool that renders from the formula instead.
- Which scientific subjects do AI image models handle well?
- Photographic and abstract ones: laboratory scenes, glassware, crystals, dissolution, flames, light trails, space imagery and non-specific molecular decoration. The common factor is that there is no specific fact available to get wrong.
- Why is text in AI images so unreliable?
- Text is rendered as visual pattern rather than as characters, so the model reproduces the look of writing without a representation of the words. Short technical strings — element symbols, units, subscripts — are the worst case, and they are exactly what scientific images need.
- Is it acceptable to use AI images in educational material?
- For decoration, yes — backgrounds, borders, characters, covers. For anything a student learns a fact from, no. A generated diagram will contain errors and will look authoritative, and the student has no way to detect the difference.