Automated Model Selection for Tabular Data

Avinash Amballa; Gayathri Akkinapalli; Manas Madine; Naga Pavana Priya Yarrabolu; Przemyslaw A. Grabowicz

Automated Model Selection for Tabular Data

Avinash Amballa, Gayathri Akkinapalli, Manas Madine, Naga Pavana Priya Yarrabolu, Przemyslaw A. Grabowicz

TL;DR

This work tackles the challenge of selecting predictive feature interactions in tabular data, where naive grid searches are computationally prohibitive. It introduces two automated approaches—Priority-based Random Grid Search and Greedy Search (Forward/Backward)—to simultaneously select base and second-order interaction features within a linear modeling framework, using one-hot encoding for categoricals. Experiments on real-world Adult data reveal limited gains from interactions, while a synthetic dataset with ground-truth interactions shows near-optimal predictive power (approximate $R^2$ of $0.996$) and favorable runtimes, highlighting a trade-off between efficiency and accuracy. Overall, the results demonstrate that explicit, interpretable interaction selection can be effectively automated for tabular data, with practical implications for transparent, efficient predictive analytics on structured datasets.

Abstract

Structured data in the form of tabular datasets contain features that are distinct and discrete, with varying individual and relative importances to the target. Combinations of one or more features may be more predictive and meaningful than simple individual feature contributions. R's mixed effect linear models library allows users to provide such interactive feature combinations in the model design. However, given many features and possible interactions to select from, model selection becomes an exponentially difficult task. We aim to automate the model selection process for predictions on tabular datasets incorporating feature interactions while keeping computational costs small. The framework includes two distinct approaches for feature selection: a Priority-based Random Grid Search and a Greedy Search method. The Priority-based approach efficiently explores feature combinations using prior probabilities to guide the search. The Greedy method builds the solution iteratively by adding or removing features based on their impact. Experiments on synthetic demonstrate the ability to effectively capture predictive feature combinations.

Automated Model Selection for Tabular Data

TL;DR

) and favorable runtimes, highlighting a trade-off between efficiency and accuracy. Overall, the results demonstrate that explicit, interpretable interaction selection can be effectively automated for tabular data, with practical implications for transparent, efficient predictive analytics on structured datasets.

Abstract

Paper Structure (18 sections, 8 figures, 1 table, 3 algorithms)

This paper contains 18 sections, 8 figures, 1 table, 3 algorithms.

Introduction
Related works
Background and hypotheses
Background
Objective and Framework
Data
US Adult Income dataset and discussion
Synthetic data generation
Techniques and methods
Computational complexity of the feature space
Priority-based random grid search
Greedy search
Results
Priority-based random grid search
Greedy Search
...and 3 more sections

Figures (8)

Figure 1: Feature importance's in the real-life Adult Income dataset.
Figure 2: Causal diagram showing the interactions between the features used to create the synthetic dataset.
Figure 3: A snippet of the generated synthetic dataset Adult2.
Figure 4: Weight assignment in Adult2 creation
Figure 5: Evolution of the best feature set over iterations of the priority search algorithm. From the figure, we observe that the algorithm reached near optimal model
...and 3 more figures

Automated Model Selection for Tabular Data

TL;DR

Abstract

Automated Model Selection for Tabular Data

Authors

TL;DR

Abstract

Table of Contents

Figures (8)