This directory contains benchmarks of structured-data-models on TabArena/BeyondArena, ScoringBench, and TALENT.
Run the commands below from the repository root:
pip install . \
"autogluon.tabular>=1.6.4b20260924,<=1.7.0" \
"tabarena[data-foundry,plot,preprocessing]>=0.1.1.dev20260924104714,<=0.2.0"Note
Weights of TabFM are distributed under the TabFM Non-Commercial License v1.0.
Review the license before running the TabFM benchmark, which will download its weights noninteractively.
-
TabICLv2:python -m benchmark.tabular.tabarena.main --model tabiclv2
-
KumoTabular:python -m benchmark.tabular.tabarena.main --model kumo-tabular-large python -m benchmark.tabular.tabarena.main --model kumo-tabular-medium python -m benchmark.tabular.tabarena.main --model kumo-tabular-small
-
TabFM:python -m benchmark.tabular.tabarena.main --model tabfm
Pass a dataset name to run only that TabArena dataset:
python -m benchmark.tabular.tabarena.main \
--model kumo-tabular-large \
--dataset blood-transfusion-service-centerEvaluate all available model results with:
python -m benchmark.tabular.tabarena.evaluate-
TabICLv2:python -m benchmark.tabular.beyondarena.main --model tabiclv2
-
KumoTabular:python -m benchmark.tabular.beyondarena.main --model kumo-tabular-large python -m benchmark.tabular.beyondarena.main --model kumo-tabular-medium python -m benchmark.tabular.beyondarena.main --model kumo-tabular-small
-
TabFM:python -m benchmark.tabular.beyondarena.main --model tabfm
By default, each command evaluates the recommended core subset.
Repeat --subset to combine filters:
python -m benchmark.tabular.beyondarena.main \
--model tabiclv2 \
--subset core \
--subset groupedPass a dataset name to run only that BeyondArena dataset.
Use --subset lite for its first split:
python -m benchmark.tabular.beyondarena.main \
--model tabiclv2 \
--dataset parkinsons_biomedical_voice_measurements \
--subset liteAvailable subset filters include problem types (classification, regression), size buckets (tiny, small, medium, large), split regimes (iid, temporal, grouped), feature groups (low-dim, high-dim, text, high-cardinality), and split selections (core, lite, all). Prefix a filter with ! to negate it.
Evaluate all available model results with:
python -m benchmark.tabular.beyondarena.evaluate