Sitemap
Towards AI

We train AI Engineers. We deploy Enterprise AI. Join 100K+ AI practitioners at Towards AI Academy or reach out for help making your company AI native. towardsai.com

Computer Vision

Object Detection — Document Layout Analysis Using Monk AI

Article on comparative study between 3 different object detection architectures- YOLOv3, Faster-RCNN, and SSD512 on the task of identifying various regions of a document and using MonkAI to load these models in just a few lines of code.

11 min readAug 13, 2020

--

Press enter or click to view image in full size
Example of Document Layout Analysis using Object Detection
Example of Document Layout Analysis using Object Detection

Introduction

This is an article on how Object Detection can help us in predicting various regions of a document. It can be useful in cropping out headlines, paragraphs, tables, images, etc. from a document image that can be later processed to get desired information from them as per the need. We compare the performance of 3 different Object Detection Architectures for this task, i.e., YOLOv3, Faster-RCNN, and SSD512, and use Monk Library to load these models.

Detailed Tutorial on Github.

About the Dataset

The training dataset used for this task is PRImA Layout Analysis Dataset. It includes a wide variety of different document types, reflecting various challenges in layout analysis. Particular emphasis is placed on:

  • Magazine scans from a variety of mainstream news, business, and technology publications which contain a mixture of simple and complex layouts (e.g. non-Manhattan, with varying font sizes, etc.)
  • Technical articles on a variety of disciplines, including papers in journals and conference proceedings, with both simple and complex layouts present.

The dataset contains 18 labels, namely, ‘caption’, ‘chart’, ‘credit’, ‘drop-capital’, ‘floating’, ‘footer’, ‘frame’, ‘graphics’, ‘header’, ‘heading’, ‘image’, ‘linedrawing’, ‘maths’, ‘noise’, ‘page-number’, ‘paragraph’, ‘separator’ and ‘table’

It can be downloaded from here.

Monk AI :

Monk object detection is a collection of all object detection pipelines. The benefit is two-fold for each pipeline- make the installation compatible for multiple OS, Cuda versions, and python versions, and make it low code with a standardized flow of things. Monk object detection enables a user to solve a computer vision problem in very few lines of code. For this task, we’ll be using 3 different pipelines of this library for 3 different architectures- yolov3, gluoncv_finetune, and mxrcnn.

Table of Contents

  1. Installing Monk Object Detection Toolkit
  2. Using the Pre-trained model for the Document Layout Analysis Task
  3. Training your own Model
  • Downloading and Pre-Processing Data (Format Conversion, Selective Data Augmentation)
  • Training the model from Scratch

4. Inference and Comparison

1. Installing Monk Object Detection Toolkit

First of all, clone the library to your system using the following command:

! git clone https://github.com/Tessellate-Imaging/Monk_Object_Detection.git

Then, choose the pipeline that you want to install and the correct requirements file of that pipeline depending on your system’s CUDA version or Colab version. These are the commands for the pipelines that I’ve used for this task:

#For yolov3 (used for yolov3 architecture)
! cd Monk_Object_Detection/7_yolov3/installation && cat requirements.txt | xargs -n 1 -L 1 pip install
#For gluoncv_finetune (used for SSD512 architecture)
! cd Monk_Object_Detection/1_gluoncv_finetune/installation && cat requirements_cuda10.1.txt | xargs -n 1 -L 1 pip install
#For mxrcnn (used for FasterRCNN architecture)
! cd Monk_Object_Detection/3_mxrcnn/installation && cat requirements_cuda10.1.txt | xargs -n 1 -L 1 pip install

For more pipelines or ways to install visit Monk Object Detection Library.

2. Using the Pre-trained model for the Document Layout Analysis Task

If you don’t want to train the model on your own, and just want to use the model that we’ve trained for the task, you can use the following piece of code to directly use it:

For YOLOv3:

import os
import sys
from IPython.display import Image
sys.path.append("Monk_Object_Detection/7_yolov3/lib")
from infer_detector import Infer
gtf = Infer()

Download and initialize the pre-trained model:

! wget --load-cookies /tmp/cookies.txt "https://docs.google.com/uc?export=download&confirm=$(wget --save-cookies /tmp/cookies.txt --keep-session-cookies --no-check-certificate 'https://docs.google.com/uc?export=download&id=1Si1puABMiijtvLvH-XMnr2pVj4K2lUkO' -O- | sed -rn 's/.*confirm=([0-9A-Za-z_]+).*/\1\n/p')&id=1Si1puABMiijtvLvH-XMnr2pVj4K2lUkO" -O obj_dla_yolov3_trained.zip && rm -rf /tmp/cookies.txt! unzip -qq obj_dla_yolov3_trained.zip! mv dla_yolov3/yolov3.cfg .f = open("dla_yolov3/classes.txt")
class_list = f.readlines()
f.close()
model_name = "yolov3"
weights = "dla_yolov3/dla_yolov3.pt"
gtf.Model(model_name, class_list, weights, use_gpu=True, input_size=416)

And you can test it:

#change test1 to whatever image you want it to test for.
img_path = "test1.jpg"
gtf.Predict(img_path, conf_thres=0.3, iou_thres=0.5)
Image(filename='output/test1.jpg')

For SSD512:

import os
import sys
sys.path.append("Monk_Object_Detection/1_gluoncv_finetune/lib/")
from inference_prototype import Infer

Download and initialize the pre-trained model:

! wget --load-cookies /tmp/cookies.txt "https://docs.google.com/uc?export=download&confirm=$(wget --save-cookies /tmp/cookies.txt --keep-session-cookies --no-check-certificate 'https://docs.google.com/uc?export=download&id=1E6T7RKGwy-v1MUxVJm-rxt5XcRyr2SQ7' -O- | sed -rn 's/.*confirm=([0-9A-Za-z_]+).*/\1\n/p')&id=1E6T7RKGwy-v1MUxVJm-rxt5XcRyr2SQ7" -O obj_dla_ssd512_trained.zip && rm -rf /tmp/cookies.txt! unzip -qq obj_dla_ssd512_trained.zipmodel_name = "ssd_512_vgg16_atrous_coco";
params_file = "dla_ssd512/dla_ssd512-vgg16.params";
class_list = ["paragraph", "heading", "credit", "footer", "drop-capital", "floating", "noise", "maths", "header", "caption", "image", "linedrawing", "graphics", "fname", "page-number", "chart", "separator", "table"]
gtf = Infer(model_name, params_file, class_list, use_gpu=True)

And you can test it:

#change test1 to whatever image you want it to test for.
img_name = "test1.jpg"
visualize = True
thresh = 0.3
output = gtf.run(img_name, visualize=visualize, thresh=thresh)

For Faster-RCNN:

import os
import sys
sys.path.append("Monk_Object_Detection/3_mxrcnn/lib/")
sys.path.append("Monk_Object_Detection/3_mxrcnn/lib/mx-rcnn")
from infer_base import *

Download and initialize the pre-trained model:

! wget --load-cookies /tmp/cookies.txt "https://docs.google.com/uc?export=download&confirm=$(wget --save-cookies /tmp/cookies.txt --keep-session-cookies --no-check-certificate 'https://docs.google.com/uc?export=download&id=1TZQSBiMDBrGhcT75AknTbofirSFXprt8' -O- | sed -rn 's/.*confirm=([0-9A-Za-z_]+).*/\1\n/p')&id=1TZQSBiMDBrGhcT75AknTbofirSFXprt8" -O obj_dla_faster_rcnn_trained.zip && rm -rf /tmp/cookies.txt! unzip -qq obj_dla_faster_rcnn_trained.zipclass_file = set_class_list("dla_fasterRCNN/classes.txt")set_model_params(model_name="vgg16", model_path="dla_fasterRCNN/dla_fasterRCNN-vgg16.params")set_hyper_params(gpus="0", batch_size=1)set_img_preproc_params(img_short_side=300, img_long_side=500, mean=(196.45086004329943, 199.09071480252155, 197.07683846968297), std=(0.25779948968052024, 0.2550292865960972, 0.2553027154941914))initialize_rpn_params()
initialize_rcnn_params()
sym = set_network()
mod = load_model(sym)

And you can test it:

#change test1 to whatever image you want it to test for.
set_output_params(vis_thresh=0.9, vis=True)
Infer("test1.jpg", mod);

3. Train Your Own Model

Data Preparation

The dataset can be downloaded using the following command:

! wget --load-cookies /tmp/cookies.txt "https://docs.google.com/uc?export=download&confirm=$(wget --quiet --save-cookies /tmp/cookies.txt --keep-session-cookies --no-check-certificate 'https://docs.google.com/uc?export=download&id=1iBfafT1WHAtKAW0a1ifLzvW5f0ytm2i_' -O- | sed -rn 's/.*confirm=([0-9A-Za-z_]+).*/\1\n/p')&id=1iBfafT1WHAtKAW0a1ifLzvW5f0ytm2i_" -O PRImA_Layout_Analysis_Dataset.zip && rm -rf /tmp/cookies.txt! unzip -qq PRImA_Layout_Analysis_Dataset.zip

All the images in the dataset are in TIFF format. Training on TIFF images was over 5x slower than JPEG format images because of their huge size. Therefore, TIFF images were converted to JPEG format images.

for name in glob.glob(root_dir+img_dir+'*.tif'):     
im = Image.open(name)
name = str(name).rstrip(".tif")
name = str(name).lstrip(root_dir)
name = str(name).lstrip(img_dir)
im.save(final_root_dir+ img_dir+ name + '.jpg', 'JPEG')

The data is present in the VOC format. To use it with various pipelines, we first convert it to Monk format, which is directly compatible with a lot of Monk pipelines, and later on, we can easily convert it to some other format if required. If you want to skip converting to Monk format and want to directly convert it to some other required format, then you can check out that pipelines’ example notebooks here.

Get Swapnil Ahlawat’s stories in your inbox

Join Medium for free to get updates from this writer.

Monk Format

./Document_Layout_Analysis/ (final_root_dir)
|
|-----------Images (img_dir)
| |
| |------------------img1.jpg
| |------------------img2.jpg
| |------------------.........(and so on)
|
|
|-----------train_labels.csv (anno_file)

Annotation file format

| Id         | Labels                                 |
| img1.jpg | x1 y1 x2 y2 label1 x1 y1 x2 y2 label2 |
  • Labels: xmin ymin xmax ymax label
  • xmin, ymin — top left corner of the bounding box
  • xmax, ymax — bottom right corner of the bounding box

The code for data conversion is straight-forward but very long. You can check out the code in one of the notebooks here.

Following are the format requirements for various pipelines used for this task:

  1. yolov3 pipeline used for YOLOv3 architecture required data in YOLOv3 format. You can check out this conversion in this notebook.
  2. gluoncv-finetune pipeline used for SSD512 architecture directly takes in Monk Format for training. So, there was no need for further conversion.
  3. mxrcnn pipeline used for Faster-RCNN architecture required data in COCO format. You can check out this conversion in this notebook.

Selective Data Augmentation

There was an issue with the dataset. As most part of a document is text, there were far more paragraphs in the dataset than there were other labels such as tables or graphs. To handle this huge bias in the dataset, we augmented only those document images which had one of these minority labels in them. For example, if the document only had paragraphs and images, then we didn’t augment it. But if it had tables, charts, graphs or any other minority label, we augmented that image by many folds. This process helped in reducing the bias in the dataset by around 25%. This selection and augmentation has been done during the format conversion from VOC to Monk Format. You can check out the code in one of the notebooks here.

For data augmentation, we have used the Albumentations library. It offers a lot of different ways to augment data, such as random cropping, translation, hue, saturation, contrast, brightness, etc. You can check more about this library here. It can be directly installed using pip command:

! pip install albumentations

Following is the function that we wrote for data augmentation. There were few cases where bounding boxes were going out of the image and Albumentations library wasn’t able to handle it, so we’ve written a custom function to make sure that labels are inside the image.

def augmentData(fname, boxes):
image = cv2.imread(final_root_dir+img_dir+fname)
image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)

transform = A.Compose([
A.IAAPerspective(p=0.7),
A.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.1, rotate_limit=5, p=0.5),
A.IAAAdditiveGaussianNoise(),
A.ChannelShuffle(),
A.RandomBrightnessContrast(),
A.RGBShift(p=0.8),
A.HueSaturationValue(p=0.8)
], bbox_params=A.BboxParams(format='pascal_voc', min_visibility=0.2))

for i in range(1, 9):
label=""
transformed = transform(image=image, bboxes=boxes)
transformed_image = transformed['image']
transformed_bboxes = transformed['bboxes']
#print(transformed_bboxes)
flag=False
for box in transformed_bboxes:
x_min, y_min, x_max, y_max, class_name = box
if(xmax<=xmin or ymax<=ymin):
flag=True
break
label+= str(int(x_min))+' '+str(int(y_min))+' '+str(int(x_max))+' '+str(int(y_max))+' '+class_name+' '

if(flag):
continue
cv2.imwrite(final_root_dir+img_dir+str(i)+fname, transformed_image)
label=label[:-1]
combined.append([str(i) + fname, label])

Calculating Mean and Standard deviation of the dataset

The mxrcnn pipeline (used for Faster-RCNN) also requires mean and standard deviation as one of the parameters. It can be calculated using the following function:

def normalize():
channel_sum = np.zeros(3)
channel_sum_squared = np.zeros(3)
num_pixels=0
count=0
for file in files:
file_path=final_root_dir+img_dir+file
img=cv2.imread(file_path)
img= img/255.
num_pixels += (img.size/3)
channel_sum += np.sum(img, axis=(0, 1))
channel_sum_squared += np.sum(np.square(img), axis=(0, 1))
mean = channel_sum / num_pixels
std = np.sqrt((channel_sum_squared/num_pixels) - mean**2)

#bgr to rgb conversion
rgb_mean = list(mean)[::-1]
rgb_std = list(std)[::-1]
return rgb_mean, rgb_std
mean, std = normalize()
mean=[x*255 for x in mean]

Train Your Own Model

This is where the real power of Monk Library kicks in. Writing code for Object detection architectures can be a very tedious task, but it can be achieved in very few lines of code using Monk Object Detection Library.

For the comparison purposes, all 3 architectures have been trained for 30 epochs with a learning rate of 0.003.

For YOLOv3:

import os
import sys
sys.path.append("Monk_Object_Detection/7_yolov3/lib")
from train_detector import Detector
gtf = Detector()
#dataset directories
img_dir = "Document_Layout_Analysis/Images/"
label_dir = "Document_Layout_Analysis/labels/"
class_list_file = "Document_Layout_Analysis/classes.txt"
gtf.set_train_dataset(img_dir, label_dir, class_list_file, batch_size=16)
gtf.set_val_dataset(img_dir, label_dir)
gtf.set_model(model_name="yolov3")
#sgd is found out to perform better than adam optimiser on this task
gtf.set_hyperparams(optimizer="sgd", lr=0.003, multi_scale=False, evolve=False)
gtf.Train(num_epochs=30)

For Faster-RCNN:

import os
import sys
sys.path.append("Monk_Object_Detection/3_mxrcnn/lib/")
sys.path.append("Monk_Object_Detection/3_mxrcnn/lib/mx-rcnn")
from train_base import *# Dataset params
root_dir = "./";
coco_dir = "Document_Layout_Analysis"
img_dir = "Images"
set_dataset_params(root_dir=root_dir, coco_dir=coco_dir, imageset=img_dir);
set_model_params(model_name="vgg16")
set_hyper_params(gpus="0", lr=0.003, lr_decay_epoch='20', epochs=30, batch_size=8)
set_output_params(log_interval=500, save_prefix="model_vgg16")
#Preprocessing image parameters(mean and std calculated during data pre-processing)
set_img_preproc_params(img_short_side=300, img_long_side=500, mean=(196.45086004329943, 199.09071480252155, 197.07683846968297), std=(0.25779948968052024, 0.2550292865960972, 0.2553027154941914))
initialize_rpn_params();
initialize_rcnn_params();
#Removing cache if any
if os.path.isdir("./cache/"):
os.system("rm -r ./cache/")
roidb = set_dataset()
sym = set_network()
train(sym, roidb)

For SSD512:

import os
import
sys
sys.path.append("Monk_Object_Detection/1_gluoncv_finetune/lib/");
from detector_prototype import Detector
gtf = Detector()
root = "Document_Layout_Analysis/"
img_dir = "Images/"
anno_file = "train_labels.csv"
batch_size=8
gtf.Dataset(root, img_dir, anno_file, batch_size=batch_size)#vgg16 architecture, with atrous convolutions, pretrained on COCO dataset is used for this task
pretrained = True
gpu=True
model_name = "ssd_512_vgg16_atrous_coco"
gtf.Model(model_name, use_pretrained=pretrained, use_gpu=gpu)
gtf.Set_Learning_Rate(0.003)
epochs=30
params_file = "saved_model.params"
gtf.Train(epochs, params_file)

These models were trained on 16GB of NVIDIA Tesla V100. YOLOv3 took the least amount of time in training- 6–7 hrs, SSD512 took around 11 hrs, and Faster-RCNN took the most amount of time- 24+ hrs.

4. Inference and Comparison

The inference code is almost the same as the one used when directly using the pre-trained model. You can check them out in the notebooks here.

Following results were obtained on test images after training the model from scratch:

Results Obtained from YOLOv3:

The outputs produced by YOLOv3 were very accurate. It’s the only model that was able to identify drop-capital among the 3 architectures. Though the confidence in the predictions is low compared to other models, their classification is most accurate among all three.

Press enter or click to view image in full size
Press enter or click to view image in full size
Press enter or click to view image in full size
Inference on Test Images from YOLOv3 Architecture

Results Obtained from Faster-RCNN:

Faster-RCNN detected bounding boxes with very high confidence, but it missed some of the important regions, such as footer in the 1st example, heading in the 2nd example, and drop capital in the 3rd. If we decrease the threshold confidence for getting the missing boxes, it produces a lot of random boxes with no clarity of what it represents.

Press enter or click to view image in full size

Results Obtained from SSD512:

SSD512 produces outputs with very high confidence, a lot of them being 0.9+. It was also the only model that was able to identify footer and noises like division lines in the document. But it was also producing repetitive or incorrect headings such as ‘floating’ in the 2nd example (extra box with incorrect label), and graphics and paragraph in the third (2 boxes with different labels for the same region).

Press enter or click to view image in full size
Inference on Test Images from SSD512 Architecture

Following inferences can be made from this tutorial on the basis of their output:

  1. Monk library makes it very easy for students, researchers and competitors to create deep learning models and try different hyper-parameter tuning to increase the accuracy of the model in very few lines of code.
  2. Faster-RCNN gave the worst performance on this task, whereas SSD512 and YOLOv3 gave comparable results.
  3. If you want to use a model which shouldn’t take much time to train and missing minute details like footers or separators won’t affect your work, go for YOLOv3.
  4. If these small details are crucial for your work and the focus is more on bounding box prediction than classification, go for SSD512. It should also be considered that gluoncv-finetune pipeline of Monk AI (which has been used for SSD512) also provides architectures that are pre-trained on various other datasets, such as COCO dataset.

References:

  1. GitHub Repository for the Tutorial: https://github.com/swapnil-ahlawat/Document_Layout_Analysis-MonkAI
  2. Dataset: https://www.primaresearch.org/dataset/
  3. Monk Object Detection Library: https://github.com/Tessellate-Imaging/Monk_Object_Detection
  4. Albumentations Library: https://github.com/albumentations-team/albumentations

Thanks for Reading! I hope you find this article informative & useful. Do share your feedback in the comment section! You can connect with me on LinkedIn.

--

--

Swapnil Ahlawat
Swapnil Ahlawat

Written by Swapnil Ahlawat

3rd Year B.E. Computer Science student at BITS Goa. Interested in Data Structure & Algorithms, Computer Vision and Web Development.

Towards AI
Towards AI

Published in Towards AI

We train AI Engineers. We deploy Enterprise AI. Join 100K+ AI practitioners at Towards AI Academy or reach out for help making your company AI native. towardsai.com