TaRL explores how large language models can be repurposed for tabular few-shot learning by reprogramming their serialization pipeline and inference procedures. This repository packages the training/evaluation helpers, model definitions, and experiment scripts that produced the accompanying paper.
tarl/: Python package with reusable components.experiments/: reproducible experiment suites (see catalog below).evaluate/: CLI entry points that execute TaRL or baseline evaluators.models/: TaRL heads and backbone wrappers.serializers/: utilities for turning tabular samples into LM-friendly sequences and embeddings.utils/: runtime helpers (caching, reporting, etc.).
config/: YAML configs for datasets, serializers, preprocessors, and models.scripts/: cluster helpers (SLURM job arrays, rsync sync, etc.).exp/: runtime outputs (created after running experiments).tabarena-datasets*: metadata dumps used by selected scripts.
-
Install Python 3.12 or use the pixi workspace
python -m venv .venv source .venv/bin/activate pip install -e .
or with pixi:
pixi install
-
Preparing datasets
- TabArena datasets can be automatically downloaded because they are available from openml.
- CARTE benchmark datasets must be downloaded manually from huggingface.
- All scripts live under
tarl/experimentsand can be launched with:python -m tarl.experiments.e01_baselines
- Each script writes configs, reports, and plots under
exp/<experiment-name>/so reruns pick up where they left off. Check the generated*_todo.txtfiles for outstanding configurations. - Pixi task shortcuts mirror the most common experiments:
pixi run e01→tarl.experiments.e01_baselinespixi run e03→tarl.experiments.e03_gamma_tuningpixi run e04→tarl.experiments.e04_autogammapixi run e08→tarl.experiments.e08_comprehensivepixi run plots→tarl.experiments.final_results
- Baseline or TaRL evaluations can also be triggered directly via YAML:
python -m tarl.evaluate.tarl exp/e01-baselines/config/<run>.yamlpython -m tarl.evaluate.baseline_fewshot exp/e01-baselines/config/<run>.yaml
-
The embedding server can be launched with:
pixi run vector-store
or
python -m tarl.utils.embedding_server
-
The data collator talks to this server to compute and cache the embeddings. Once all the embeddings necessary are cached, the server does not need to stay running any more.
- Utility & Debug
config.py: shared helpers for naming conventions, config emission, and report gathering.
- Baselines & Kernel Studies
e01_baselines.py: main CARTE classification benchmark with LM ablations and baselines.e02_kernel_sim.py: inspects intra-/inter-class embedding similarities from cached tensors.e05_serialization.py: compares serialization strategies by regenerating datasets with alternative encodings.e06_runtime.py: remeasures TaRL runtime under batching/caching variants.e09_knn_embedding.py: evaluates k-NN baselines on LM-derived row embeddings.
- Gamma Analysis & Selection
e03_gamma_tuning.py: sweeps gamma values for TaRL inference and summarizes best settings.e03_1_gamma_analysis.py: correlates tuned gammas with meta-features and visualizes backbone trends.e04_autogamma.py: evaluates automatic gamma scaling heuristics.e08_1_comprehensive_gamma.py: reprocesses comprehensive sweeps with alternative gamma selectors.
- Meta-Learning & Comprehensive Runs
e08_comprehensive.py: end-to-end benchmark spanning CARTE/TabArena tasks, including sweeps and baselines.e08_2_comprehensive_meta.py: scales gamma meta-learning with richer feature sets and per-dataset holdouts.e08_3_comprehensive_meta_binary.py: frames gamma selection as binary classification per candidate value.
- Aggregation & Reporting
final_results.py: stitches experiment outputs into publication-ready plots and tables.meta_features.py: current meta-feature engineering utilities (seemeta_features_old.pyfor the previous iteration).
Run any script with python -m tarl.experiments.<module>; each maintains its
own config directory and CSV summaries inside exp/<module-name>/.
- Experiment outputs live under
exp/<experiment-name>/and include:config/: generated YAML configs for reruns.tables/&plots/: aggregated metrics and visualizations.*_todo.txt: outstanding runs to launch.- this directory should not be modified manually -- the scripts should handle all updates.
- Evaluator reports are cached under
DATA_DIR/models/<run>/alongside logits and metadata. scripts/houses helper shell scripts (SLURM job arrays, remote sync, etc.) you can adapt to your infrastructure.
- Pre-download backbone checkpoints to avoid repeated Hugging Face downloads.
- Post-processing notebooks expect CSVs emitted by the experiment scripts; avoid editing
exp/artifacts by hand to keep runs reproducible.
This work was supported by IBM through the IBM-Rensselaer Future of Computing Research Collaboration.
@inproceedings{kangLanguageModelRepresentations2026,
title = {Language {{Model Representations}} for {{Efficient Few-Shot Tabular Classification}}},
booktitle = {Proceedings of the {{ACM Web Conference}} 2026},
author = {Kang, Inwon and Ram, Parikshit and Zhou, Yi and Samulowitz, Horst and Seneviratne, Oshani},
year = 2026,
pages = {4069--4080},
doi = {10.1145/3774904.3792477},
}