✓ Indexed in BIBNEX
Info:eu Repo/semantics/article
Deep Learning Based Text Extraction from Video Using CNN, LSTM, and Transformer Models
Abstract
This study offers a deep learning-based method for text extraction from video frames, addressing issues like motion blur, variable text orientations, and background noise. Traditional optical character recognition (OCR) methods like Tesseract suffer from these problems, while contemporary deep learning models offer notable advancements. The suggested model uses Convolutional Neural Networks (CNNs) to identify text regions, Transformer-based models to increase recognition accuracy, and Long Short-Term Memory (LSTM) networks to maintain sequences. Several tests demonstrate that by striking a balance between accuracy and real-time functionality, the CNN + LSTM architecture performs better than conventional OCR algorithms. The results show that transformer-based methods have the highest accuracy but the highest computational cost. deep learning models like CNN, LSTM, and Transformers can handle contextual recognition, temporal sequencing, and spatial detection, they are particularly well suited for video text extraction. This hybrid approach, in contrast to traditional OCR, guarantees high accuracy even in video frames that are noisy, blurry, or multilingual.
Keywords
Optical Character Recognition
Text Extraction
Deep Learning
Convolutional Neural Networks
Long Short-Term Memory
Cite This Article
(2025).
Deep Learning Based Text Extraction from Video Using CNN, LSTM, and Transformer Models.
Journal of Recent Innovations in Computer Science and Technology
, 2(3)
.
https://doi.org/10.70454/jricst.2025.20304
“Deep Learning Based Text Extraction from Video Using CNN, LSTM, and Transformer Models.”
Journal of Recent Innovations in Computer Science and Technology,
vol. 2,
no. 3,
2025
.
https://doi.org/10.70454/jricst.2025.20304
“Deep Learning Based Text Extraction from Video Using CNN, LSTM, and Transformer Models.”
Journal of Recent Innovations in Computer Science and Technology
2
, no. 3
(2025)
.
https://doi.org/10.70454/jricst.2025.20304
(2025)
‘Deep Learning Based Text Extraction from Video Using CNN, LSTM, and Transformer Models’,
Journal of Recent Innovations in Computer Science and Technology
, 2
(3)
.
Available at:
https://doi.org/10.70454/jricst.2025.20304
Deep Learning Based Text Extraction from Video Using CNN, LSTM, and Transformer Models.
Journal of Recent Innovations in Computer Science and Technology.
2025
;2
(3)
.
doi:
10.70454/jricst.2025.20304