Models that have to work on real hardware, not just in a notebook: object detection that hands cutting coordinates to a robot cell, a convolutional classifier built up milestone by milestone, and a decision tree you can poke at live in the browser.
Dismantling lithium-ion cells for recycling is normally done with sensors and actuators — slow, error-prone, and wasteful. The brief was to replace that with a machine-learning vision system that finds the exact lines a robot should cut along, in three stages of the teardown.
YOLOv11Oriented bounding boxesRaspberry Pi 5TCP / PLCOpenCV
Focus 1 — housingA cut line placed just under the top edge of the metal casing, 3–5 mm clear of the terminals, with its two end points printed in pixel coordinates.Focus 2 — electrodesCopper and aluminium foils told apart by class and confidence, then separated by oriented cut lines — as few cuts as the geometry allows.Focus 3 — tape, end viewAdhesive strips holding the wound layers together are located on the roll and given short cut segments.Focus 3 — tape, rotatingThe same detector tracking tape as the cell is turned by hand, so the cuts follow the strip rather than the graphite underneath.
Live inference on a Raspberry Pi 5 — green overlays are the proposed cut lines, magenta text the coordinates sent on to the controller.
The three focuses
Focus 1 — battery housing. Find a cut line just below the top edge of the metal casing so the terminals and electrodes can be sawn off without damaging what is inside. The camera sits 350–400 mm above the cell and the line should land 3–5 mm from the terminals.
Focus 2 — electrodes. Detect the copper (anode) and aluminium (cathode) sections and define cut lines that separate them in as few pieces as possible.
Focus 3 — adhesive tape. Locate the tape strips holding the layers together, even when they are partly hidden, and return cuts that do not reach the graphite underneath.
The success criteria set for the work were a detection error inside 1 mm, robustness to different backgrounds and tape positions, and a working link into the plant's PLC.
How it runs
Everything runs on a Raspberry Pi 5 with a Pi camera. A TCP server accepts two text commands from a PLC or PC — inference runs detection on a fresh image, coordinates returns the cut-line points — which we exercised with the Hercules TCP client before the plant link existed. Live video inference scripts were delivered alongside it. YOLO model sizes n through x were compared; the practical difference was inference time rather than accuracy.
Models and tooling
Ultralytics YOLOv11, trained 100 epochs from pre-trained weights on an Nvidia A100 in Google Colab. Focus 1 uses the medium model with ordinary upright boxes — several objects, orientation irrelevant. Focuses 2 and 3 use YOLOv11m-obb, the rotated-box variant, because there the cut line is an orientation.
Datasets were annotated in Roboflow with five classes: big face, terminal, copper, aluminium and adhesive tape. Post-processing is NumPy vector maths on the box corners; OpenCV and Matplotlib draw the overlays; Picamera grabs frames and a plain socket server carries the results.
Where it falls short
Sensitive to lighting and texture changes, which occasionally flips the material identified in Focus 2.
Overlapping or partly hidden tapes remain hard.
It detects geometry, not defects or flaws.
It was never tested against the actual robot — only the TCP interface it would speak to.
The recommendations that came out of it: faster hardware such as an Nvidia Jetson, and a more varied dataset.
Raspberry Pi 5, Pi Camera 2, router/switch, customer PLC
Interface
TCP/IP to a Siemens PLC — inference and coordinates commands
02 / Deep learning
A squeeze-and-excitation CNN that labels hand signs.
A sign-language interpreter has a large but quirky labelled archive and a small handful of labelled examples from a second, different-looking camera — plus 7,029 images nobody has labelled at all. The task: train a classifier on the archive that still works on the new images, and auto-label the rest.
Four convolutional stages, two of them channel-gated by squeeze-and-excitation, then a two-layer classifier head.
The squeeze-and-excitation block, after Hu, Shen & Sun (CVPR 2018) — the channels each input actually needs get turned up, the rest get turned down.
HandSignCNN — channel attention and forward passPyTorch
# squeeze-and-excitation: re-weight channels by their own global statisticsclassSEBlock(nn.Module):
def__init__(self, channels: int, reduction: int = 8):
self.avg_pool = nn.AdaptiveAvgPool2d(1) # H×W → 1×1
self.fc = nn.Sequential(
nn.Linear(channels, channels // reduction, bias=False), # bottleneck
nn.ReLU(inplace=True),
nn.Linear(channels // reduction, channels, bias=False), # restore
nn.Sigmoid(), # 0…1 gates
)
defforward(self, x):
b, c, _, _ = x.size()
y = self.avg_pool(x).view(b, c) # squeeze → (B, C)
y = self.fc(y).view(b, c, 1, 1) # excite → (B, C, 1, 1)return x * y # scale → broadcast multiplyclassHandSignCNN(nn.Module):
defforward(self, x):
x = F.relu(self.bn1(self.conv1(x))) # (1,32,32) → (32,32,32)
x = F.relu(self.bn2(self.conv2(x))) # (32,32,32) → (64,32,32)
x = self.pool1(x) # → (64,16,16)
x = F.relu(self.bn3(self.conv3(x))) # (64,16,16) → (128,16,16)
x = self.se3(x) # channel attention
x = self.pool2(x) # → (128,8,8)
x = F.relu(self.bn4(self.conv4(x))) # (128,8,8) → (256,8,8)
x = self.se4(x) # channel attention
x = self.pool3(x) # → (256,4,4)
x = self.dropout_features(x) # feature-level dropout
x = x.view(x.size(0), -1)
x = F.relu(self.fc1(x)) # → (B, 256)
x = self.dropout_fc(x)
return self.fc2(x) # → (B, 25) logits
Concepts used
Channel attentionSqueeze-and-excitation blocks on stages 3 and 4. Widening the bottleneck (reduction 16 → 8) was the last gain: 92% → 93.1%.
Kaiming initFan-out Kaiming-normal on conv and linear layers, zero biases, BatchNorm at γ=1/β=0 — chosen because the network is all ReLU, and it stabilised a suspiciously fast convergence.
Average over max poolingSwapping max pooling for average pooling kept more of the weak stroke detail in 32×32 grayscale hands.
Two-level dropoutp = 0.5 on the feature map and again before the classifier, added when train/val accuracy ran far ahead of test accuracy.
Domain-aware augmentationLight affine + jitter on the big archive; a stronger version on the 143 target-domain examples. No flips — a mirrored hand sign is a different letter.
Fine-tuning on the target distributionTraining on the archive alone stalled; fine-tuning on the small augmented target set closed the domain gap and jumped accuracy by ~31 points.
Mixed precisionautocast + GradScaler for the training and fine-tuning loops.
Stratified split80/20 train/validation, stratified by label, with the validation half deliberately left un-augmented.
How it got to 93.1%
≈ 50% — three conv layers with batch norm and max pooling, tuned by hand. Hyperparameter sweeps stopped helping.
81.6% — strongly augmenting the 143 labelled target examples and fine-tuning on them: the training archive and the test images simply did not look alike.
90% — a fourth convolution and dropout layers to kill the overfitting that fine-tuning exposed.
90.5% — Kaiming initialisation in place of PyTorch's default, for stabler convergence.
92% — squeeze-and-excitation blocks, and average pooling instead of max.
93.1% — lr 3e-4, weight decay 1e-4, 40 epochs, and an SE reduction ratio of 8 rather than 16.
A last finding worth keeping: identical seeds still moved the score by 1–2 points, because the layers were constructed straight onto the CUDA device while torch.manual_seed had only seeded the host. Passing the device at layer construction made runs comparable again.
Data
27,455 labelled archive images — three channels, but only one carried signal, and each image randomly red, green or blue. Converted to single-channel grayscale at 32×32.
143 labelled examples from the target set, plus 7,029 unlabelled images to predict.
Everything normalised to mean 0.5, std 0.5; predictions written out as CSV.
Framework
PyTorch ≥ 2.x, CUDA, mixed precision
Input
1 × 32 × 32 grayscale, normalised 0.5 / 0.5
Optimiser
Adam, lr 3e-4, weight decay 1e-4
Batch / epochs
256 · 40 epochs, then 150 fine-tuning epochs
Loss
Cross-entropy (label smoothing tried and dropped)
Context
Introduction to Deep Learning, THWS — with Dev Patel
03 / Live app
Above $65 an hour? A decision tree on IBM HR data.
A small interactive classifier: set the attributes of an employee and watch the tree decide whether they earn above or below $65 an hour, with the path through the tree shown as it goes.