Final Year Thesis
CNN-Based Bangla Handwritten Character Recognition: Exploring Ekush Dataset for Performance Enhancement
Benchmarking 16 Deep Learning & Classical Machine Learning Models on Isolated Characters
Abstract
Developed an optical character recognition (OCR) pipeline to accurately classify complex handwritten Bangla characters (vowels and consonants). Evaluated and benchmarked 16 machine learning and deep learning models on a curated subset of the Ekush dataset containing 80,000 images across 50 classes. DenseNet achieved peak state-of-the-art recognition accuracy (97.93% on vowels, 95.99% on consonants) by maximizing feature reuse and gradient flow under compute-constrained training.
Problem
Bengali Handwritten Character Recognition (BHCR) poses major challenges due to intricate cursive strokes, horizontal connecting lines (Matra), structural similarities across consonants, and high handwriting variability. Furthermore, prior research focused heavily on numerals, leaving isolated character recognition under-explored on large, diverse datasets under constrained compute environments.
Key metrics
- VOWEL ACCURACY
- 97.93%
- CONSONANT ACCURACY
- 95.99%
- DATASET VOLUME
- 80,000
- MODELS BENCHMARKED
- 16 Classifiers
Methodology
The research implemented an empirical comparative framework spanning 16 algorithmic architectures. Deep learning models (DenseNet, Custom 2-Layer CNN, and LeNet-5) processed 2D spatial feature representations, while classical machine learning classifiers (SVM, Extra Trees, Random Forest, Bagging, KNN, Logistic Regression, SGD, Nearest Centroid, Naive Bayes) were trained on normalized tabular pixel vectors.
Data pipeline
- Grayscale conversion and intensity normalization ([0, 1] range)
- Spatial uniform image resizing to 28×28 and 32×32 pixel matrices
- Tabular pixel matrix flattening for classical machine learning algorithm ingestion
- Stratified class partitioning ensuring balanced representation across all 50 character classes
Benchmarks
| Model | Vowel | Consonant |
|---|---|---|
| DenseNet | 97.93% | 95.99% |
| Custom 2-Layer CNN | 92.96% | 87.72% |
| LeNet-5 | 91.63% | 85.35% |
| Support Vector Machine (SVM) | 88.14% | 78.77% |
| Extra Trees Classifier | 85.05% | 72.15% |
| Random Forest | 83.64% | 71.6% |
| Bagging Classifier | 81.2% | 69.45% |
| K-Nearest Neighbors (KNN) | 79.5% | 67.8% |
| Decision Trees | 76.4% | 64.3% |
| Logistic Regression | 75.1% | 63.85% |
| Stochastic Gradient Descent (SGD) | 73.9% | 62.1% |
| Nearest Centroid | 68.4% | 56.9% |
| Gaussian Naive Bayes | 62.7% | 51.3% |
Takeaways
- DenseNet achieved the highest performance across all metrics (97.93% vowel accuracy, 95.99% consonant accuracy) by leveraging tight interlayer connections for gradient propagation and feature reuse.
- Non-linear Support Vector Machines (88.14% vowel, 78.77% consonant) outperformed all other tree-based and distance-based classical algorithms, establishing the best non-neural baseline.
- Consonant recognition presented a significantly higher error rate across all 16 classifiers due to shared structural strokes, subtle ligature differences, and higher intra-class handwriting variability compared to vowels.
- Feature reuse in densely connected layers allowed superior accuracy within just 5 training epochs, demonstrating high efficiency under constrained computational resources.