Abstract
Tabular data is the most prevalent form of structured data, necessitating robust models for classification and regression tasks. Traditional models like eXtreme Gradient Boosting (XGBoost) have gained popularity for their strong performance, while deep learning models such as Tabular Retrieval-Augmented Generation (TabR) and TabNet offer innovative approaches. TabR uniquely employs Retrieval-Augmented Generation (RAG) to reduce uncertainty and enhance predictive accuracy, whereas TabNet relies on a sequential attention mechanism without incorporating RAG. This study systematically compares TabR and TabNet in classification and regression tasks using benchmark datasets, with evaluations based on accuracy, Area Under the Curve (AUC), Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R2 (Coefficient of Determination). The results reveal that TabR, with its RAG component, outperforms XGBoost in classification, effectively managing uncertainty. However, in regression tasks, XGBoost continues to excel over TabR. Meanwhile, TabNet performs comparably but lacks the performance enhancement provided by RAG in TabR. These findings highlight the potential of RAG in deep learning models for tabular data classification and suggest areas for further exploration in improving regression performance.
| Original language | English |
|---|---|
| Pages (from-to) | 191719-191732 |
| Number of pages | 14 |
| Journal | IEEE Access |
| Volume | 12 |
| DOIs | |
| Publication status | Published - 2024 |
Keywords
- Deep learning models
- equilibrium
- machine learning
- retrieval-augmented generation (RAG)
- TabNet
- TabR
- tabular data
- XGBoost
Fingerprint
Dive into the research topics of 'Tabular Data Classification and Regression: XGBoost or Deep Learning with Retrieval-Augmented Generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver