TY - GEN
T1 - AI-Based Book Classification Using Book Titles
T2 - 11th IEEE International Conference on Computing, Engineering and Design, ICCED 2025
AU - Fahmi, Faisal
AU - Sofiyani, Zulfatun
AU - Margono, Hendro
AU - Gunarti, Endang
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Assigning library classification codes improves shelving and browse-ability but remains time-consuming and inconsistent across institutions. This paper presents an AI-based method to predict Library of Congress Classification (LCC) classes and Subclasses from book titles only by using Project Gutenberg records. The study compares single- and two-stage prediction, where Stage-1 predicts the first letter of the LCC (AZ) (e.g., P for Language and Literature) and Stage-2 ranks twoletter LCC classes (e.g., PR and PS) within that first-letter group. We benchmark a short-text baseline (TF-IDF with a logistic regression) against a fine-tuned transformer (DistilBERT) under author-grouped train/validation/test split sets. Experiments report Top-1/Top-3 accuracy and Macro-F1, analyze class imbalance under Full and Merge-Rare label schemes, and include a focused case study on adjacent two-letter LCC classes to characterize short-title confusions. Results show that title-only classification is feasible, where the single-stage DistilBERT model achieves the strongest Macro-F1 and the twostage variant improves the baseline (especially under the Full scheme) by reducing off-class errors, but decreases the effectiveness of DistilBERT. Besides, Top-3 suggestions support human cataloging when only titles are available.
AB - Assigning library classification codes improves shelving and browse-ability but remains time-consuming and inconsistent across institutions. This paper presents an AI-based method to predict Library of Congress Classification (LCC) classes and Subclasses from book titles only by using Project Gutenberg records. The study compares single- and two-stage prediction, where Stage-1 predicts the first letter of the LCC (AZ) (e.g., P for Language and Literature) and Stage-2 ranks twoletter LCC classes (e.g., PR and PS) within that first-letter group. We benchmark a short-text baseline (TF-IDF with a logistic regression) against a fine-tuned transformer (DistilBERT) under author-grouped train/validation/test split sets. Experiments report Top-1/Top-3 accuracy and Macro-F1, analyze class imbalance under Full and Merge-Rare label schemes, and include a focused case study on adjacent two-letter LCC classes to characterize short-title confusions. Results show that title-only classification is feasible, where the single-stage DistilBERT model achieves the strongest Macro-F1 and the twostage variant improves the baseline (especially under the Full scheme) by reducing off-class errors, but decreases the effectiveness of DistilBERT. Besides, Top-3 suggestions support human cataloging when only titles are available.
KW - Book Title
KW - Library of Congress Classification
KW - Quality Education
UR - https://www.scopus.com/pages/publications/105033054401
U2 - 10.1109/ICCED68324.2025.11324715
DO - 10.1109/ICCED68324.2025.11324715
M3 - Conference contribution
AN - SCOPUS:105033054401
T3 - 2025 IEEE 11th International Conference on Computing, Engineering and Design, ICCED 2025 - Proceedings
BT - 2025 IEEE 11th International Conference on Computing, Engineering and Design, ICCED 2025 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 13 November 2025 through 15 November 2025
ER -