Multimodal Cancer Classification Challenge

Classifying cells from oral brush samples as coming from cancer patients or healthy ones, with only 19 patients. We placed ninth, and it taught me not to trust a validation score.

Technologies Used

PyTorchResNetDenseNet backbonescross-attention fusion
Multi-modal cancer classification

The problem

The task was to predict whether a single cell image comes from a cancer patient or a healthy one. The cells come from oral brush scrapes, stained and imaged two ways, brightfield and fluorescence, as 128 by 128 crops. There are only 19 patients, and every cell carries its patient's diagnosis as its label. The idea is that cells near a tumour change subtly, but it also means plenty of cells labelled as cancer look completely normal. So the real question wasn't how well a model fits the data, it was how well it works on a patient it has never seen.

What we did

The first thing we had to get right was the split. If you split cells at random, the same patient ends up in both training and validation, and the AUC looks great for the wrong reason. So we split by patient, which left a tiny validation set that swung depending on which patients we held out. I also wanted to know whether a CNN could simply learn who the patient is, so I trained one to predict patient ID among 12 patients. It did far better than chance, which told me that identity is strongly encoded in these images, and heavier augmentation only reduced it partly.

Architecture
Architecture

For the model we used a two-stream ResNet, with one stream for each image type, fused before a small classifier. We tried ResNet-18, 34 and 50, DenseNet-121, and a few ways of fusing the streams: sum, concatenation, gating and attention. The backbone stayed frozen for the first half of training and was then fine-tuned. On top of that we used early stopping, augmentation, dropout, class weighting and test-time augmentation.

The result

Our best model, a fine-tuned two-stream ResNet with multiscale cross-attention, reached an AUC of 0.9999 on training and 0.9135 on validation, but only 0.7818 on the public test set. We placed ninth, and the top score was 0.8651. Fine-tuning beat a frozen backbone and test-time augmentation helped, while ensembling folds and training for longer didn't.

Comparison Table
Comparison Table
Results
Results

What went wrong

The gap between 0.91 and 0.78 says it all: with so few patients, my validation score wasn't something I could trust. I also spent too much time swapping backbones when I should have been thinking about where to fuse the two images. The group that topped the challenge used early fusion, which combines the two images before the network, while our fusion happened after the two backbones. And I didn't write down what I tried, so I couldn't even learn from it properly. That's why I document failures now.

---

When: May 2026 Stack: PyTorch, ResNet and DenseNet backbones, cross-attention fusion Role: team of 5