Convolutional neural networks were the architecture that proved deep learning could solve real-world vision (AlexNet, 2012). For a decade, they were the default for any image task. Vision Transformers (ViTs) and their descendants have largely overtaken them on benchmarks since 2020, but CNNs remain the right pick for several specific cases — and the inductive biases they exploit are still relevant whether or not the parameter count is a transformer.
A convolution slides a small filter (kernel) across the input, computing a weighted sum at each position. Stacking convolutions builds increasingly abstract feature detectors:
Two mechanical tricks make this work:
These inductive biases match the structure of natural images. For tasks where they don't (text, tabular data), CNNs lose to architectures with different priors.
A canonical CNN classifier:
Image (224×224×3)
→ Conv (3×3, 64 filters), ReLU, BatchNorm
→ Conv (3×3, 64 filters), ReLU, BatchNorm
→ MaxPool (2×2) [output: 112×112×64]
→ Conv (3×3, 128 filters), ReLU, BatchNorm
→ Conv (3×3, 128 filters), ReLU, BatchNorm
→ MaxPool [output: 56×56×128]
→ ...
→ Global average pool [output: 1024]
→ Dense (num_classes)
→ Softmax
The pattern: progressively halve spatial dimensions; double channel dimensions; conv-conv-pool repeated. End with global pooling and a classifier.
This is the recipe behind VGG, ResNet, EfficientNet, and most production vision networks before transformers.
ResNet introduced residual connections — a skip connection that adds the input of a block to its output. Allowed training of much deeper networks (152 layers and beyond) without vanishing gradients.
output = layer(x) + x # the residual connection
The math is the same as a regular layer, but the gradient flows back through both the layer and the skip path. Identity mapping prevents gradient decay.
By 2026, residual connections appear in nearly every deep network — CNN or transformer. ResNet-50 remains a credible baseline for narrow vision tasks.
For new vision projects in 2026:
Vision transformers have surpassed CNNs on most benchmarks, but the calculus is more nuanced for production:
CNNs are usually 5-10× more efficient than transformers at the same task on small images. Mobile CPUs and dedicated NPUs (Apple Neural Engine, Snapdragon) are often optimised for convolutions specifically.
For on-device tasks (face detection, photo segmentation, scanner OCR), CNNs remain the production default.
Transformers typically need more data to train than CNNs. With < 10k labelled images, fine-tuning a pretrained CNN often beats training a ViT from scratch.
That said, fine-tuning a pretrained ViT (CLIP, DINOv2) on small data is competitive — the pretrained backbone covers the data hunger.
Some vision tasks are genuinely about local features: defect detection on a manufacturing line, OCR character recognition, simple medical imaging. CNN inductive biases match the task; no need for global attention; CNNs are faster.
Some segmentation and detection benchmarks still see CNNs (or hybrid CNN-transformer architectures) winning. The U-Net family for medical imaging, YOLOv8/v9 for real-time detection.
CNNs as feature extractors plus task-specific heads:
Modern pipelines often replace the CNN backbone with a transformer (Swin, ViT). The task heads (detection, segmentation) remain similar.
CNN-specific training discipline:
For deploying CNNs:
A typical deployment chain: train in PyTorch → export to ONNX → optimise with TensorRT → serve. Each step gives maybe 1.5-3× speedup.
For a typical "I need to recognise / detect / segment something":
This produces a competitive baseline in days instead of weeks. Push for transformer backbones only when you have data and compute.