✓ Indexed in BIBNEX
Info:eu Repo/semantics/article
Feature Selection for Automatic Answer Extraction from Online Web Forums
Abstract
This research proposes an efficient method for automatic answer quality classification in online web forums using a Multi-Task BERT-Based deep learning framework. The primary objective is to accurately categorize user responses into low, medium, and high-quality classes by leveraging advanced language representation and relevant content features. Starting with Stack Overflow forum data collection, the methodology moves on to comprehensive preprocessing, exploratory data analysis, and the extraction of syntactic and semantic features. Syntactic features include things like sentence length, punctuation use, code snippets, and TF-IDF vectors and contextual embeddings using spaCy. To reduce redundancy and increase the quality of the model input, we utilised feature selection with RFECV and the principle component analysis (PCA). A two-output-layer multi-task BERT architecture, which the model uses as its basis, can handle major classification and auxiliary tasks simultaneously, allowing it to increase generalisation. After seven training epochs, the trained model achieved remarkable results: 99.54% Training Accuracy, 92.13% Validation accuracy, 92.13% Precision, 92.13% Recall, and 92.13% F1-Score. In addition to providing fast inference, the model typically only takes 0.0071 milliseconds per sample for reactions. In large-scale community-driven QA platforms, these results validate the model's robustness, efficiency, and appropriateness for real-time applications.
Keywords
Answer Quality Classification
Multi-Task BERT
Feature Selection
Online Web Forums and Deep Learning
Citation
Feature Selection for Automatic Answer Extraction from Online Web Forums.
Journal of Recent Innovation in Science and Technology
.
2026.
Vol. 2
(1)
DOI: 10.70454/jrist.020101