Skip to content
AI for Beginners

Computer Vision

The area of AI that lets computers make sense of images and video: recognising objects, reading text in a photo, spotting faces, or describing a scene. If natural language processing is AI learning to read and write, computer vision is AI learning to see. It is what lets a phone sort your photos or a car notice a pedestrian.

Computer vision is the branch of AI focused on visual information. Its job is to take raw pixels, which mean nothing on their own, and turn them into something useful: this is a dog, that is a road sign, these words appear on the page, this face matches that one.

You meet it more often than you might think. It powers the way your phone groups photos by person, the check-in gates that read passports, the tools that turn a photographed receipt into text, and the systems that help a car sense what is around it. Medical scans, factory quality checks, and wildlife monitoring all lean on it too.

Increasingly, computer vision is being folded into general AI tools. When you can paste a photo into a chatbot and ask what is in it, or point your camera at a menu in another language and get a translation, you are using computer vision working alongside language ability. AI that combines senses like this is called multimodal AI.

The word is handy mainly for understanding what a tool can do. If a product mentions computer vision, you can expect it to work with pictures or video in some way, whether that is recognising, sorting, reading, or describing what it sees.

Related terms