Sitemap
Artificial Intelligence in Plain English

New AI, ML and Data Science articles every day. Follow to join our 3.5M+ monthly readers.

Top 6 most common mistakes in Deep Learning pipeline

In this article we’ll see how simple practices when applied together can significantly improve model accuracy, also easing out and simplify the training process.

11 min readDec 17, 2020

--

Article Overview

I’ve seen a lot of DL projects where people would be using the latest and the greatest from what deep learning has to offer, to finish up a project, and repeating some of the very basic mistakes again and again. And I’ve been there too, but ever since I’ve started working for a startup, it made me responsible enough to not be lazy anymore, and not overlook the basics. So let’s have a glance at some of the silly mistakes I used to make, that I can bet you most likely do too.

1. Training model from scratch

So unless you’re only implementing a project to learn something new, you should never ever even think about train a model from scratch.

I’m kind of stretching a bit too much here, but still I’d stick to it. I know that training a model from scratch sounds cool and fun, because you’d think doing this would prove your capabilities, that you know so much that you were able to do yada yada yada, but actually this is just a terrible mistake.

Now there’s nothing wrong with training a model from scratch, I mean all those pretrained models were trained from scratch itself in the first place. The problem is, using a pretrained model, at least as one of the building blocks of your project-specific model, it would not only result in significantly better accuracy but you’d save a ton of time, especially when the model has to interpret the input data having a lot of features.

Press enter or click to view image in full size
source

A few months back I was working on a project based on video classification. It was a very simple and straight forward task. And at that time Pytorch didn’t have support for pretrained models for video classification, so I trained a 3D ResNet model from scratch, with over 30 million parameters, and it took me over a week to achieve the base-level accuracy.

Luckily, with the next release after a few months, Pytorch started supporting 3 pretrained models for video classification, and guess what, I just casually tried to see how a pretrained model would train and perform as compared to my scratch one, and you’d not believe it, within a day, a single goddam day, I was able to outperform 1 week of hassle with over a 15% increase in accuracy, and that with only half the parameters. I mean it was like I was using a holy new technology for my project.

So, next time have mercy on you and your project with the blessings of a pretrained model.

2. Incorrectly training a pretrained model

(i) Not applying required preprocessing

Okay, so you decided to do transfer learning, cool, but guess what? You didn’t check how the training data for that pretrained model was preprocessed, and you are still using your own preprocessing.

Every fancy DL framework right now, at least to my knowledge, that provides pretrained models also provides the exact preprocessing needed to be applied to the data before passing it to the model.

And this is really important, you might think that since this pretrained model was trained on another domain of data, let’s say images of birds, and what you are dealing with is totally different, let’s say CT scans, so since the domain is different thus you can change the preprocessing too, wrong.

The model doesn’t care what you’re dealing with, all it cares is what are the numbers its being given. If the new data doesn’t have the same statistics as previous data, then the gradients of the very first blocks of the model with respect to the inputs would be too high, plus at the beginning of the training the loss would be very high too, so right after few batch iterations the weights will change too much, and the model will forget a lot of whatever feature maps it learned from previous training. Thus killing the actual purpose of transfer learning.

So, next time when using a pretrained model, make sure to look at the documentation and find out what preprocessing is required to be applied to the input data.

(ii) Wrong transfer learning

So you have applied the required preprocessing too, now what? Well, firstly you would replace the last few layers with new ones such that the total number of output units are equal to the number of classes you have, okay, all good till here too.

But guess what? You didn’t freeze the rest of your model apart from the newly replaced layers. The thing is, since the last of the few layers are initialized with totally random parameters, hence the gradients of the loss with respect to the parameters of those layers would be too high. And there’s nothing wrong with that actually, this is exactly what’s supposed to happen.

But since the gradients with respect to those activations would also be used to compute gradients with respect to the rest of the parameters, so you’ll be applying too much updates to those parameters. And again this time as well, the pretrained model would just forget whatever it learned from previous training.

So, next time make sure to freeze the unreplaced part of the pretrained model before starting the training.

(iii) Not fine-tuning

Alright, you’ve chosen your pretrained your model, replace the last few layers, frozen the rest, and trained the model with good accuracy. Once again, guess what? Your not yet utilizing the fullest of what that pretrained model has to offer.

This mistake is quite easy to do, because right after you’re done training the model, you would get pretty good results, thanks to the pretrained model. But your whole model is still under fitted because you never trained previously frozen parameters.

So here’s the road map:

  • Firstly, freeze the unreplaced layers and train the model.
  • After enough training when the loss is quite low, you unfreeze those layers and continue the training.

Until now the loss is low, hence you wouldn’t be pushing too much updates to those parameters.

The idea is, you’d simply just tune those parameters ever so slightly, such that the model would just become more better without forgetting what it has learned from previous training.

Hence the name fine-tuning.

3. Finishing up project too early

I honestly used to do this all the time. I know performing a lot of experimentation is both time consuming and boring.

Press enter or click to view image in full size
Actually its worth being bored |Photo by Ethan Hu on Unsplash

Different experiments if done correctly can provide a lot of improvements for both performance and accuracy. Now sometimes you just don’t have enough time to do enough experiments due to some tight schedule. So what should you experiment the most in that limited time?

Get Rishik C. Mourya’s stories in your inbox

Join Medium for free to get updates from this writer.

I think optimizing hyperparameters in this case, wouldn’t be a good idea. Yes, you would be able to select the best hyperparams after the experimentations, but since per experiment, you’d want more visible changes to the results and hyperparameter optimization is not gonna give you much of a day and night difference.

So in my opinion, testing out different model architectures would be our best bet. You don’t have to fully train the model up to its full potential, instead, train different architectures for time that makes the ratio (total parameters / training seconds) same for each experiment. With this you could compare how well and fast each architecture can perform. And my suggestion would be to at least try 5 to 8 architectures.

If your training data is too much, then per experiment don’t use all of it to speed things up, just take a fraction of it and continue, and don’t forget to use same fractional data every experiment.

4. Ignoring the basics

In ML pipeline there’s just so much to take care of, hence most of the times it becomes just a force of habit to forget the basics. Okay, so this again might be boring if you’d think, but basic techniques like regularizing parameters, applying clipping to the gradients, using learning rate schedulers etc, can really polish your final results especially when you’re using a massive model with respect to your data.

The issue is, a simple and smaller model would always be under fitted to the data no matter how hard you train. And there is no such rule of thumb that we can use to decide the parameters counts which would just perfectly align to the specific dimension of the data you have.

So now you’d make the model bigger and better, and along that comes the overfitting. Dropout is good, I mean really good, but it’s not like that you can apply dropout to all the layers, otherwise extremely slow training would occur for obvious reasons. Thus for bigger models, using regularization techniques is your best bet.

Now when using bigger models, you’d be giving a free pass to vanishing gradients. And utilizing residual or dense connections between different layers is, of course, better than not using them. But they both come along with some more issues, for instance, if you have limited CUDA memory then you might not be able to train big models with residual or dense connections.

This is the reason why gradient clipping is soo damn cool. It does not cost any kind of performance or accuracy regression, in fact, it makes both of them better.

Okay then, since your model is big so is your weights space. The idea behind weight space is that it’s a huge multidimensional space, with the dimensions equal to the number of parameters of your model.

Now with very high dimensional weight space, comes more number of local minima. Thus it becomes an obvious choice to utilize some technique that enables our model to explore all local minima, and finally, settle at the global one, or at least, at the minimum of all local minima that model has explored yet.

This is the part where the family of cyclic learning rate schedulers comes in. The idea is to increase and decrease the learning rate in a cyclic manner, thus allowing the model to discover many minima, and once the model is deep down to a convex region that is really deep, then in the next cycle the model would most likely stay in that convex region. And thus settling the model.

Disclaimer: It’s not that these scheduler 100% guarantee to settle your model to the global minima, obviously, but yes they work really really great almost all the time. And these scheduler has become one of the de-facto choices for getting state of the art results.

5. Ignoring data bias

This is one of the most common things to ignore. Most of the time we do take care of limited data by applying data augmentations, ignoring the additional bias that we have created. Since, for the class that has limited training data, even though you have applied augmentations, but they would still share too much of the same spatial relations, for instance, the texture in the image, RGB histogram, etc. Random horizontal flipping, rotation can’t provide enough augmentation to such relations.

So applying simple augmentations here actually creates even more data bias in this case.

The solution to this problem is not easy, but you should at least try not to ignore this. Try getting more data somehow if possible by any means. Tons of free datasets are available on the internet, at least some of them might correlate to what you have.

You can write your own web scraping script that extracts the data for you. Or better still, if possible then at least try creating yourself.

One solution that works okay most of the cases, is to use class weights.

The idea is to normalize the updates for each class so that each parameter would be updated giving equal weightage to all the classes, even though you’d have data bias in each batch.

So before computing the mean of the loss per batch, you would multiply the loss for each output unit by a scalar value. And for each class, that scalar value is computed by (1 / num of examples in this class) * (total number of examples) / number of classes.

Scaling by (total number of examples / number of classes) helps keep the loss to a similar magnitude, so the sum of the class weights of all examples stays the same.

6. Only using validation and test data to get satisfied

Training the model using training data, validating on validation data, and testing different models on testing data to choose the best model, is not at all a complete solution. Yes, you heard me right.

That is just the least you can do, but you should keep in mind where this model that you’ve been training for days would be deployed exactly, and what kind of data it would have to deal with, just try thinking all the possible cases. And thus try modifying the model or the preprocessing, such that the model is not too much dependent on the specificity of the data you have.

It is of course not that you would be able to cover 100% of all the data that your model might see during inference, but just try covering as much of scenarios as possible.

You must have seen this situation when the model was doing good when given familiar data, but giving completely horrible results when provided with confusion.

For instance, recently the news came out, where an object detection model gets dizzy, and mistakes ref’s bald head for a soccer ball.

source: twitter

There is actually nothing wrong with the model. If you look at the head and the ball yourself, I can bet even you wouldn’t be able to conclude which one is which. Below I’ve cropped down the head and the ball, could you tell them apart?

Press enter or click to view image in full size
Left one is ball and right one is the head.

This is a perfect example when we get satisfied with the results using testing data (that might even contain biased data), ignoring the actual new data that might come during inference.

It’s not easy to do much about it, but we should at least take note of the cases that might cause our model to be failed. And try fixing these issues with new data or some other practices.

Conclusion

Getting a noticeable jump in the model’s improvement is not that hard. Following basics and applying them together could make any model significantly better in most of the cases. It is just the fact that sometimes we overlook the simpler solutions because they don’t sound cool to us anymore. We should be okay with a simple linear regression if it solves our problem. With that said, I hope you’ve realized how much you could have improved your previous projects you’ve done so far.

Alright, thanks for reading and I hope to see you again.

--

--