Skip to main navigation Skip to search Skip to main content

Evaluating TabPFN: a transformer-based foundation model for explainable health insurance claim prediction

Research output: Contribution to journalArticlepeer-review

Abstract

Background: Predictive modeling for companies involved in health insurance policies remains an active area of actuarial research as they seek to leverage Machine Learning (ML) applications to increase productivity and improve operational efficiency. Despite the success of classical machine learning models on tabular insurance datasets, they often require feature engineering and hyperparameter tuning, hence posing scalability and deployment challenges. While addressing these limitations and an identified gap in the literature, this research evaluated a next-generation Transformer-based tabular foundation model, the Tabular Prior-Data Fitted Network (TabPFN-2.5, Prior Labs, Germany), for health insurance claim prediction and benchmarked its performance against established ML models. Methods: The study utilizes a primary dataset comprising 13,904 records with 13 features, along with three additional validation datasets to assess robustness across heterogeneous data distributions. Our approach involves extensive preprocessing, including data cleaning and converting categorical features into a numerical data type. Model performance was extensively evaluated using metrics of Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and R-squared (R2). Results: The results showed that the TabPFN regression model consistently outperforms the baseline models, achieving an R2 of 0.9776 on the primary dataset and demonstrating better generalization across validation datasets. The impact of each feature on the model's prediction was evaluated using SHapley Additive exPlanations (SHAP) as an Explainable Artificial Intelligence (XAI) method, revealing smoking status, age, and body mass index as the most influential determinants. Conclusion: This research highlights the impact of TabPFN on improving predictive modeling of healthcare finances and on selecting the most suitable policies for customers.

Original languageEnglish
Pages (from-to)1880390
Number of pages1
JournalFrontiers in Public Health
Volume14
DOIs
StatePublished - 2026

Bibliographical note

Publisher Copyright:
Copyright © 2026 Shawosh, Nisar and Alwahaishi.

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 1 - No Poverty
    SDG 1 No Poverty

Keywords

  • explainable AI
  • foundation model
  • health insurance
  • machine learning
  • regression

ASJC Scopus subject areas

  • Public Health, Environmental and Occupational Health

Fingerprint

Dive into the research topics of 'Evaluating TabPFN: a transformer-based foundation model for explainable health insurance claim prediction'. Together they form a unique fingerprint.

Cite this