Final Year Thesis

CNN-Based Bangla Handwritten Character Recognition: Exploring Ekush Dataset for Performance Enhancement

Benchmarking 16 Deep Learning & Classical Machine Learning Models on Isolated Characters

Author
Marcel David Baroi
Institution
Daffodil International University
Project
#26095
Period
2023 — 2024

Abstract

Developed an optical character recognition (OCR) pipeline to accurately classify complex handwritten Bangla characters (vowels and consonants). Evaluated and benchmarked 16 machine learning and deep learning models on a curated subset of the Ekush dataset containing 80,000 images across 50 classes. DenseNet achieved peak state-of-the-art recognition accuracy (97.93% on vowels, 95.99% on consonants) by maximizing feature reuse and gradient flow under compute-constrained training.

Problem

Bengali Handwritten Character Recognition (BHCR) poses major challenges due to intricate cursive strokes, horizontal connecting lines (Matra), structural similarities across consonants, and high handwriting variability. Furthermore, prior research focused heavily on numerals, leaving isolated character recognition under-explored on large, diverse datasets under constrained compute environments.

Key metrics

VOWEL ACCURACY
97.93%
CONSONANT ACCURACY
95.99%
DATASET VOLUME
80,000
MODELS BENCHMARKED
16 Classifiers

Methodology

The research implemented an empirical comparative framework spanning 16 algorithmic architectures. Deep learning models (DenseNet, Custom 2-Layer CNN, and LeNet-5) processed 2D spatial feature representations, while classical machine learning classifiers (SVM, Extra Trees, Random Forest, Bagging, KNN, Logistic Regression, SGD, Nearest Centroid, Naive Bayes) were trained on normalized tabular pixel vectors.

Data pipeline

  • Grayscale conversion and intensity normalization ([0, 1] range)
  • Spatial uniform image resizing to 28×28 and 32×32 pixel matrices
  • Tabular pixel matrix flattening for classical machine learning algorithm ingestion
  • Stratified class partitioning ensuring balanced representation across all 50 character classes

Benchmarks

ModelVowelConsonant
DenseNet97.93%95.99%
Custom 2-Layer CNN92.96%87.72%
LeNet-591.63%85.35%
Support Vector Machine (SVM)88.14%78.77%
Extra Trees Classifier85.05%72.15%
Random Forest83.64%71.6%
Bagging Classifier81.2%69.45%
K-Nearest Neighbors (KNN)79.5%67.8%
Decision Trees76.4%64.3%
Logistic Regression75.1%63.85%
Stochastic Gradient Descent (SGD)73.9%62.1%
Nearest Centroid68.4%56.9%
Gaussian Naive Bayes62.7%51.3%

Takeaways

  • DenseNet achieved the highest performance across all metrics (97.93% vowel accuracy, 95.99% consonant accuracy) by leveraging tight interlayer connections for gradient propagation and feature reuse.
  • Non-linear Support Vector Machines (88.14% vowel, 78.77% consonant) outperformed all other tree-based and distance-based classical algorithms, establishing the best non-neural baseline.
  • Consonant recognition presented a significantly higher error rate across all 16 classifiers due to shared structural strokes, subtle ligature differences, and higher intra-class handwriting variability compared to vowels.
  • Feature reuse in densely connected layers allowed superior accuracy within just 5 training epochs, demonstrating high efficiency under constrained computational resources.
← School