Recognizing Digits in Noisy CAPTCHAs Without Deep Learning

A 4-stage classical image-analysis pipeline (FFT denoising, HOG features, SVM) that reads digits out of noisy, stripe-patterned CAPTCHA images — no deep learning or OCR allowed, 92.2% test accuracy.

Technologies Used

MATLABHOG FeaturesSVMFFTImage Segmentation
Captcha

Overview

This was the mini-project for Introduction to Image Analysis (1MD110) at Uppsala University, building on five earlier assignments from the course. The task: given a CAPTCHA image showing three or four printed digits (only ever 3, 4, or 5) buried under a diagonal stripe noise pattern, recognize them — with the constraint that we couldn't use OCR or deep neural networks, only classical image analysis and traditional ML. Worked on it with a project partner.

Approach

We landed on a 4-stage pipeline: clean the image, predict how many digits it has, segment it into individual digits, then classify each one. Cleaning turned out to be the hardest part — the noise wasn't random, it was a strong diagonal stripe pattern sitting at specific frequencies, so we used an FFT notch filter to knock those frequencies out, then a distance-transform step to strip away whatever thin noise lines survived that. For recognition we used HOG features with SVM classifiers rather than anything deep-learning based, including a specialist model just for telling 3 and 5 apart, since the main classifier kept confusing those two.

Screenshot 2026-09-07 at 2.46.45 PM
Screenshot 2026-09-07 at 2.46.45 PM
Screenshot 2026-09-07 at 2.46.57 PM
Screenshot 2026-09-07 at 2.46.57 PM
Screenshot 2026-09-07 at 2.47.05 PM
Screenshot 2026-09-07 at 2.47.05 PM

What didn't work first

Our first attempt just split the image into four equal-width slots and classified whatever landed in each one. It got 86.5% on the training set but only 56.4% on the held-out test set, which told us the segmentation itself was the actual weak point, not the classifier. That's what pushed us toward proper frequency-domain denoising and the multi-stage pipeline instead of a quick fix.

Results

The final version hit 98.5% training accuracy and 92.2% on the official test set, processing each image in about 0.05 seconds — well inside the 6-minute budget for the whole set.