Visual Challenges

Intermediate

Common challenges and limitations when working with vision-enabled AI models.

Last updated: Sep 13, 2026

Common Visual Challenges

Vision systems can fail at counting, spatial relations or exact text. Resolution, training, preprocessing and task design all matter. Evaluate a specified model on reference images instead of assuming one universal failure rate.

Check the image

Generated reference image, visually checked on 2026-09-13. Answers are annotations of this image. This exercise does not run language models or simulate model answers.

How many screw heads are visible?

Enlarging a crop can make existing detail accessible. It cannot recover pixels already lost in a low-resolution input. A model comparison needs the image, question, model version and original responses, checked against this same reference.

ViT: patch projection

🔢

Counting Objects

Models often struggle to accurately count objects in images, especially when there are many similar items.

Why This Happens

A count requires identifying separate instances and avoiding omissions or duplicates. Occlusion, similar objects and task training can affect that process. Patch-based input alone does not prove a counting limit.

Common Failures

  • •Counting people in a dense crowd; measure errors for the chosen dataset and model.
  • •Counting items in a grid or array
  • •Distinguishing between "few" and "many" when items overlap

Workarounds

For critical counting tasks, consider using specialized object detection models (YOLO, Faster R-CNN) or asking the model to identify and describe each item individually rather than providing a total count.

📍

Spatial Reasoning

Understanding precise spatial relationships between objects (left/right, above/below) can be unreliable.

Why This Happens

The task needs a clear coordinate frame and precise grounding. Resizing, ambiguous viewpoints and model errors can change the answer; position embeddings alone do not guarantee exact geometry.

Common Failures

  • •Confusing left/right relationships in mirrored or symmetric images
  • •Misjudging relative distances ("closer to" or "farther from")
  • •Difficulty with rotated or unusual orientations

Workarounds

Be explicit in your prompts about which reference frame to use. Consider annotating images with visual markers or grids for critical spatial tasks.

🔤

Small Text Recognition

Fine text in images may be misread or missed entirely, especially at low resolutions.

Why This Happens

Resizing can discard strokes before the model sees them. In ViT, a learned projection processes the flattened patch pixels; details smaller than a patch are not automatically erased. Legibility also depends on contrast, training and the task.

Common Failures

  • •Misreading license plates, street signs, or small labels
  • •Confusing similar characters (0/O, 1/l/I, 5/S)
  • •Missing text in busy or low-contrast backgrounds

Workarounds

Use high-resolution images and zoom in on text regions. For critical OCR tasks, use dedicated OCR tools (Tesseract, Google Vision API, Amazon Textract) alongside or instead of vision LLMs.

👻

Visual Hallucination

Models may describe objects or details that aren't actually present in the image.

Why This Happens

Ambiguous evidence and learned language patterns can produce plausible but unsupported descriptions. This is an observable failure mode, not a complete explanation of every visual hallucination.

Common Failures

  • •Adding objects that "should" be in a scene (a keyboard near a monitor)
  • •Describing brand names or text that isn't visible
  • •Inventing details when asked about unclear regions

Workarounds

Ask the model to express uncertainty. Use prompts like "describe only what you can clearly see" or "if you cannot determine X, say so." Cross-reference critical details.

🔍

Fine Detail Recognition

Subtle details, textures, or small distinguishing features are often missed or misidentified.

Why This Happens

A patch projection is a learned transformation, not an average color. Fine details may be lost during resizing, encoding or later processing; inspect the actual input and measure the task result.

Common Failures

  • •Distinguishing between similar objects (dog breeds, car models)
  • •Reading gauges, meters, or instrument displays
  • •Identifying subtle damage or defects in inspection tasks

Workarounds

Choose a resolution that preserves the needed details within the model's input limits. Crop and focus on specific regions of interest. For specialized tasks, consider fine-tuned models trained on domain-specific data.

🖼️

Multi-Image Reasoning

Comparing multiple images requires matching objects, viewpoints or changes across inputs. Test this separately from describing one image.

Why This Happens

Cross-image tasks require correspondence: which object, time or region in one image matches another? The encoding and comparison mechanism depend on the architecture.

Common Failures

  • •Finding differences between two similar images ("spot the difference")
  • •Tracking object identity across frames
  • •Comparing fine details between product images

Workarounds

Describe each image separately first, then ask for comparison. Consider combining images into a single composite for direct comparison.

Key Takeaways

  • 1Patch size is not a universal lower bound on recognizable detail.
  • 2Evaluate counting, text and geometry separately for the chosen model and images.
  • 3A plausible visual description can contain unsupported details.
  • 4Use higher resolution, cropped regions, and explicit prompts to improve accuracy
  • 5For critical tasks, combine vision LLMs with specialized tools (OCR, object detection)
  • 6Always verify important visual information through other means