Guide home/Research directions
中文
Research directions

Computer vision: understanding images

What is in a photo, where is it, which pixels belong to an object, and how can images reveal a 3D scene? Vision makes an inviting entry point because you can see what a model gets right and wrong.

Computer vision extracts useful information from images and video. Organizing photos by person, counting road traffic, and outlining a lesion all turn pixels into decisions. Classification asks what is present, detection adds locations, and segmentation labels pixels. 3D vision also studies distance, shape, and spatial relationships.

Choose an output first, then follow how the model produces it. Convolutional networks reuse small filters across locations; vision Transformers split images into patches and let them exchange information. Architecture is only part of the story: annotation quality and differences between training and test images often determine whether a model is useful.

Research tasks
  • Recognition and retrieval: classify an image or find similar content in a collection.
  • Localization and measurement: find positions and outlines, count objects, or estimate size.
  • Scene understanding: use video or multiple views to estimate motion, depth, and 3D structure.

Clarify the input, labels, and evaluation metric before comparing networks.

An image through a vision network
  1. 01Input image
  2. 02Feature extraction
  3. 03Object detection
  4. 04Segmentation
  5. 05Depth estimation

Research milestones

  1. 2012

    Learning visual features

    AlexNet demonstrated the strength of convolutional networks on large-scale image classification. Data, GPU training, and architecture belong in the same explanation.

  2. 2016

    Making deeper networks trainable

    ResNet uses additive shortcut connections to improve optimization in deep networks. The resulting features can also support tasks beyond classification.

  3. 2016

    A separate task: recognition plus location

    YOLO predicts boxes and classes from the full image in a unified prediction problem. Detection is a different task, not a replacement for classification.

  4. 2021

    Images as sequences of patches

    ViT feeds image patches into a Transformer and studies the effect of large-scale pretraining. Convolution and attention offer design choices with different data and compute tradeoffs.

Key concepts

Convolution / CNN
A small filter moves across an image, reusing its parameters at different positions. Stacked layers combine local patterns into more complex features.
Feature / representation
Intermediate numbers that retain information useful for a task. They can feed a classifier or support detection and retrieval.
Patch and attention
A patch is a small image region. Attention uses the input to determine how much information patches exchange. A patch need not correspond to an object.
IoU and mAP
IoU measures overlap between predicted and reference boxes. mAP summarizes precision–recall performance across classes; its overlap thresholds and evaluation protocol must match when comparing scores.

Start with classification, detection and segmentation

Start with MNIST digit recognition or YOLO object detection. The former takes you through training, prediction, and a Kaggle submission; the latter begins with detections on your own photos, then training. Follow with the vision-task, training, and convolutional-network sections of CS231n.

Classification asks what, detection also asks where, and segmentation assigns meaning to individual pixels. Choose one task and examine its data and metrics.

Features, receptive fields and generalization

ResNet shows how an architectural change can be supported experimentally. For receptive fields, read Distill's visual explanation.

Inspect a prediction: what cases go wrong, does hiding the background change it, and what happens with another data source? Such observations lead to data, representations, and generalization.

Vision connects to language, generation and action

Connect images and language through multimodal models, make images through generation, and connect vision to action through embodied AI. For runtime and memory, see systems.

Vision courses and research groups

Courses and projects are in the vision catalog. For connections to robot tasks, start with SVL, RAIL, and other labs.

Representative papers

Use AlexNet and ResNet to understand feature learning and training, then choose YOLO or ViT. Follow each paper's inputs, outputs, and experiments.

2012 · NIPS 2012

ImageNet Classification with Deep Convolutional Neural Networks

ImageNet Classification with Deep Convolutional Neural Networks · Krizhevsky et al. · Fig. 2
Figure excerpt from the paper · Krizhevsky et al. · Fig. 2 · Original paper

The problem

Can a network learn recognition features from a large image collection?

The key idea

Stack convolutional layers and combine GPU training with ReLU activations and regularization.

Why this paper

It connects advances in data, compute, and training practice.

Where to start

Start with the architecture and training choices, then inspect classification results and errors.

A question to keep asking

It classifies predefined categories; it does not directly locate objects.

2016 · CVPR 2016

Deep Residual Learning for Image Recognition

Deep Residual Learning for Image Recognition · He et al. · Fig. 5
Figure excerpt from the paper · He et al. · Fig. 5 · Original paper

The problem

Why can a deeper network fit even its training data worse?

The key idea

Learn a change F(x) and add the input back, producing x + F(x).

Why this paper

A clear architectural idea is tested against plain networks at different depths.

Where to start

Read the degradation problem and residual block, then compare training errors.

A question to keep asking

Easier optimization does not make arbitrary depth beneficial; data and compute still matter.

2016 · CVPR 2016

You Only Look Once: Unified, Real-Time Object Detection

You Only Look Once: Unified, Real-Time Object Detection · Redmon et al., YOLO, Figure 2
Figure excerpt from the paper · Redmon et al., YOLO, Figure 2 · Original paper

The problem

Can one network pass predict both object categories and locations?

The key idea

Predict boxes and scores from the full image, training recognition and localization together.

Why this paper

It connects directly to a detection demo and its misses, false positives, and runtime.

Where to start

Start with the detection pipeline and grid formulation, then localization errors.

A question to keep asking

The original struggles with small or crowded objects and precise localization. Modern YOLO tools have changed substantially.

2021 · ICLR 2021

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · Dosovitskiy et al. · Fig. 1
Figure excerpt from the paper · Dosovitskiy et al. · Fig. 1 · Original paper

The problem

Does image classification require convolutional structure?

The key idea

Turn image patches into vectors and exchange information using a Transformer.

Why this paper

It connects vision to sequence processing and highlights pretraining scale.

Where to start

Follow patches into the model, then compare pretraining dataset sizes.

A question to keep asking

The original strong results rely on large-scale pretraining; they do not imply Transformers always beat CNNs on small datasets.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗