A Flexible Supervised Term-Weighting Technique and its Application to Variable Extraction and Information Retrieval
Open Access
- 11 February 2019
- journal article
- research article
- Published by IBERAMIA: Sociedad Iberoamericana de Inteligencia Artificial in INTELIGENCIA ARTIFICIAL
- Vol. 22 (63), 61-80
- https://doi.org/10.4114/intartif.vol22iss63pp61-80
Abstract
Successful modeling and prediction depend on effective methods for the extraction of domain-relevant variables. This paper proposes a methodology for identifying domain-specific terms. The proposed methodology relies on a collection of documents labeled as relevant or irrelevant to the domain under analysis. Based on the labeled document collection, we propose a supervised technique that weights terms based on their descriptive and discriminating power. Finally, the descriptive and discriminating values are combined into a general measure that, through the use of an adjustable parameter, allows to independently favor different aspects of retrieval such as maximizing precision or recall, or achieving a balance between both of them. The proposed technique is applied to the economic domain and is empirically evaluated through a human-subject experiment involving experts and non-experts in Economy. It is also evaluated as a term-weighting technique for query-term selection showing promising results. We finally illustrate the applicability of the proposed technique to address diverse problems such as building prediction models, supporting knowledge modeling, and achieving total recall.Keywords
This publication has 11 references indexed in Scilit:
- A probabilistic model derived term weighting scheme for text classificationPattern Recognition Letters, 2018
- Mining for Topics to Suggest Knowledge Model ExtensionsACM Transactions on Knowledge Discovery From Data, 2016
- Turning from TF-IDF to TF-IGM for term weighting in text classificationExpert Systems with Applications, 2016
- Evaluation and analysis of term scoring methods for term extractionInformation Retrieval Journal, 2016
- Experience-based support for human-centered knowledge modelingKnowledge-Based Systems, 2014
- A study of supervised term weighting scheme for sentiment analysisExpert Systems with Applications, 2014
- A semi-supervised incremental algorithm to automatically formulate topical queriesInformation Sciences, 2009
- Imbalanced text classification: A term weighting approachExpert Systems with Applications, 2009
- Supervised and Traditional Term Weighting Methods for Automatic Text CategorizationIEEE Transactions on Pattern Analysis and Machine Intelligence, 2008
- RANDOM WALK TERM WEIGHTING FOR IMPROVED TEXT CLASSIFICATIONInternational Journal of Semantic Computing, 2007