Run Amazon's Mitra-v2 against your gradient-boosting pipeline this month on randomly sampled data that fits its context: it is Apache 2.0 and beats tuned, ensembled XGBoost on TabArena by roughly 420 Elo. That context is 16,384 rows for binary classification and 32,768 for multiclass and regression, and the licence costs nothing. Keep XGBoost or CatBoost in production for large tables, time-ordered data and CPU-only, millisecond scoring, where the independent evidence still favours trees. Pay Prior Labs for TabPFN-3.5 only when a table outgrows Mitra's context window and you still need foundation-model accuracy, because its open weights cannot be used in production and its commercial price is not published.
The loser is XGBoost as an unexamined default. Among the tree libraries on the same board it ranks last, below CatBoost and LightGBM, and below every foundation model on that board.
| Mitra-v2 (Amazon) | TabPFN-3.5 (Prior Labs) | XGBoost (tuned) | |
|---|---|---|---|
| How it learns | In-context, plus a short fine-tune in the benchmark recipe | In-context, one forward pass | Trains a new model per dataset |
| Licence | Apache 2.0 weights and code | Weights non-commercial; production needs an enterprise licence | Apache 2.0 |
| Price (checked 6 Oct 2026) | Free | Commercial tiers: contact sales | Free |
| Row envelope | Context capped at 32,768 rows (16,384 binary) | Up to 1,000,000 rows | No ceiling |
| Features | Pretrained on 1-50 features | Up to 20,000 (6,000 recommended) | No ceiling |
| TabArena Elo | 1,774.6 (Amazon's report) | 1,866 base, 1,910 Thinking (Prior Labs' report) | 1,352.2 tuned + ensembled |
| Hardware | CUDA GPU | GPU, or Prior Labs' API | CPU is fine |
| Pick it for | Small and medium IID tables you own outright | Wide or large tables, with a budget | Large, temporal, latency-bound scoring |
| Skip it if | Your table is much bigger than the context cap | You cannot sign a commercial licence | Your IID table fits Mitra's context and you have no reason to stay |
Elo figures come from each vendor's own technical report on different leaderboard snapshots, so do not subtract one from another. Sources follow in each section.
What Changes When the Model Learns In Context
A tabular foundation model is a transformer pretrained on millions of synthetic tables that predicts labels for a new table by reading the labelled rows as context, instead of fitting a fresh model to that table. XGBoost does the opposite: every dataset gets its own trees, its own hyperparameter search and its own retraining schedule.
That difference moves cost around more than it removes it. A gradient-boosted tree is expensive to tune and almost free to serve. A foundation model skips tuning and pays at prediction time, because every prediction re-reads the context table on a GPU.
For a decade the tree won on accuracy too. Grinsztajn, Oyallon and Varoquaux at NeurIPS 2022 found tree ensembles beat neural networks on medium-sized data across 45 datasets. Pretraining on synthetic tables is what closed that gap on small data. The TabArena paper now states that foundation models "excel on smaller datasets," and that deep learning caught up with gradient-boosted trees only "under larger time budgets with ensembling."
To compare like for like, this page uses one workload throughout: a binary churn or credit-risk classifier trained on 25,000 labelled rows with 60 numeric and categorical columns, randomly split, scored as a nightly batch and through an online API with a 50 ms budget. That is the shape of most risk scorecards and per-segment churn models, and it sits inside every candidate's envelope except where noted.
Mitra-v2: The Default Challenger, With a Context Cap
Mitra-v2 is the model to test first, because it is the strongest foundation model you can deploy without a licence negotiation.
Amazon's team published the Mitra-v2 technical report on 3 September 2026: a 77-million-parameter, 2D transformer trained only on synthetic data, with weights, inference code, fine-tuning code and evaluation results released under Apache 2.0. The classifier weights on Hugging Face work with AutoGluon 1.6 and later as a drop-in replacement for the first Mitra.
The headline number: on the combined TabArena board of 51 datasets, Mitra-v2 scores 1,774.6 Elo against 1,637.5 for TabPFN-3, 1,568.9 for tuned and ensembled RealTabPFN-2.5, 1,404.0 for LightGBM, 1,394.1 for CatBoost and 1,352.2 for XGBoost, the last three tuned and ensembled. That is a vendor's own report. The mitigating fact is that TabArena's maintainers re-ran Mitra-v2 themselves among the verified models updated on 17 September, and its changelog records the Elo ratings as unchanged within their confidence intervals.
Read the limits before you plan around it, because the report states them plainly:
- Context rows are capped at 32,768 for multiclass and regression and 16,384 for binary classification. Our 25,000-row binary workload already exceeds the binary cap, so Mitra will not read all of it as context.
- Raising that cap made accuracy worse. In its negative results, the report says "the model cannot use support sets much longer than the ones it was fine-tuned on," and a few datasets got "much worse."
- The leaderboard number comes from a recipe of 50 fine-tuning steps and eight-fold bagging. The report says that system "spends more compute per dataset than forward-pass foundation models," with a median of about 8.8 minutes per classification dataset, and the Hugging Face card adds that "AutoGluon's stock defaults differ from the mitra-finetune recipe used for the reported benchmark numbers."
- The model card says the fine-tuning recipe needs a CUDA GPU.
The first Mitra, which Amazon introduced in July 2025 with state-of-the-art claims on TabRepo, TabZilla, AMLB and TabArena, was smaller still: the AutoGluon documentation lists it at 10,000 rows, 500 features and 10 classes. Those four benchmarks were run by Amazon's authors, which is why the TabArena re-run matters more than the list.
Mitra-v2 is the wrong pick for teams scoring on CPU only, teams whose training sets are hundreds of thousands of rows (it will subsample, and its own authors found longer context hurt), and anyone who expects AutoGluon's out-of-the-box settings to reproduce the 1,774.6.
TabPFN-3.5: The Broadest Envelope, Behind a Sales Call
TabPFN-3.5 is the most capable model here on paper, and the one you are least free to use.
Prior Labs' release notes put its envelope at up to 1,000,000 rows "subject to feature count and checkpoint/API limits" and up to 20,000 features, with 6,000 recommended. Its technical report gives TabPFN-3.5-Thinking 1,910 Elo, ranked 1 of 89, and base TabPFN-3.5 1,866, with a 96% win rate for Thinking against the best non-foundation model. Prior Labs ran that evaluation on one RTX PRO 6000 Blackwell.
The licence decides most purchases. The TabPFN-3.5 weights ship under tabpfn-3-5-license-v1.0, which says "the model, its derivatives, and its outputs cannot be used for any commercial or production purpose." It permits "testing, evaluation, and internal benchmarking," but it lists "competitive benchmarking for procurement" among the forbidden uses. If your team plans to run a bake-off on production data to justify a purchase, have counsel read that clause first.
On Prior Labs' pricing page, checked 6 October 2026, the API has Free, Pro and Max tiers; Pro and Max are contact-sales. Private Cloud and On-Premise each have a free non-commercial tier and a commercial tier, also contact-sales. No rate is published anywhere on the page. The release notes do disclose one pricing mechanic: successful KV-cache reuse is charged 75% less on the API. Prior Labs is also owned by SAP, which completed the acquisition on 17 July, so your negotiation is with a subsidiary of a platform vendor.
TabPFN-3.5 is the wrong pick for any team that cannot get a commercial contract signed this budget cycle, anyone who needs a list price to file a business case, and teams whose tables already fit inside Mitra-v2's context cap, where an Apache-licensed model is close enough to test first.
XGBoost and CatBoost: Where the Tree Still Wins
The tree loses the small-table leaderboard and keeps three jobs that a foundation model does badly today.
XGBoost and CatBoost are both Apache 2.0, both train on CPU or GPU, and CatBoost handles categorical columns natively. Neither has a row ceiling.
The first is large and non-IID data. The strongest evidence comes from TabArena's own maintainers. Their June 2026 paper, Beyond IID, evaluated 11 models on 142 curated datasets and found that "traditional tree-based and deep learning models still dominate on non-IID, large, and high-dimensional datasets," while foundation models do well mainly on small, IID data. Prior Labs concedes the same in its own report: "tuned and ensembled MLPs retain the highest performance on grouped, temporal, and large datasets." A churn model scored on next quarter's customers is temporal data, and a random split will flatter the foundation model.
The second is latency on CPU. A November 2025 hardware benchmark measured tree ensembles finishing inference in under 0.4 seconds, against up to 47 seconds and 4.4 GB of VRAM for TabPFN, with accuracy gains under one percentage point and no significant difference (p=0.74). That study tested the original TabPFN-1.0, so the absolute gap is out of date. The direction is the part that holds: a tree scores a row in microseconds on hardware you already run.
The third is regulated pricing work. A May 2026 study on motor insurance pricing found TabPFN "does not consistently outperform established baselines, exhibits substantially longer inference times, and is sensitive to the size of the in-context training set," and concluded it is not yet "a viable replacement for established actuarial methods."
If you stay with trees, stop defaulting to XGBoost. On the Mitra-v2 report's TabArena snapshot, tuned and ensembled LightGBM (1,404.0) and CatBoost (1,394.1) both rank above XGBoost (1,352.2). The gaps are small, but none of them favours XGBoost.
XGBoost is the wrong pick for a team whose IID table fits inside Mitra-v2's context cap and has no tuning budget: it leaves the most accuracy on the table of anything in this comparison, and it costs a hyperparameter search to get there.
What It Costs to Serve the Same Workload
On our 25,000-row workload, the tree is the cheapest to serve and the most expensive to build, and neither foundation-model vendor publishes a cost per prediction.
- XGBoost or CatBoost needs a tuning run, then serves on CPU in microseconds per row. The nightly batch and the 50 ms endpoint both fit on existing infrastructure.
- Mitra-v2 needs no tuning search, but the benchmark recipe fine-tunes and bags eight models per dataset, taking a median of about 8.8 minutes per classification dataset in a report whose experiments ran on H100 and H200 GPUs. That recipe runs on a CUDA GPU, so budget a GPU endpoint for the 50 ms path and measure it.
- TabPFN-3.5 is either a GPU you run (under a commercial licence) or the API at an unpublished rate, discounted for cache reuse. The release notes say TabPFN-3.5-Fast gives "up to 6× faster inference than the base model," and the report notes the base model runs up to 2× slower than TabPFN-3 on large training sets.
That leaves one number you have to produce yourself: GPU cost per 1,000 predictions at your real volume, with and without cached context, next to the CPU cost of the tree you already run.
How to Decide Without Regretting It
Four facts about your table predict the answer better than any leaderboard.
- Rows. IID and inside Mitra-v2's context cap (16,384 binary, 32,768 otherwise): test Mitra-v2 first. Between that and 1,000,000 with a budget: TabPFN-3.5 is the only foundation model here that fits natively. Above that: trees.
- Split type. If production data arrives later in time than training data, evaluate on a time-based split. That single change is where the Beyond IID results say trees take the lead back.
- Serving path. A CPU-only scoring service with a single-digit-millisecond budget keeps the tree, whatever the accuracy gap.
- Licence reach. If legal will not sign a non-commercial evaluation licence that bars procurement benchmarking, TabPFN-3.5 drops out before the bake-off starts.
What changes the answer: a maintainer-verified TabArena run of TabPFN-3.5-Thinking, a published Prior Labs price list, or a Mitra release with a longer usable context. Any one of those reopens the comparison.
This Week:
- List every production tabular model with its row count, column count, class count and whether its data is time-ordered. Only IID tables inside Mitra-v2's context cap (16,384 rows binary, 32,768 otherwise) go to step 3.
- Send the TabPFN-3.5 licence to counsel with one question: does an evaluation on production data that informs a purchase count as "competitive benchmarking for procurement"?
This Month:
- On two in-scope models, run Mitra-v2 through AutoGluon 1.6 (both stock defaults and the published fine-tune recipe), against your tuned CatBoost or XGBoost, on a time-based split.
- Measure GPU cost per 1,000 predictions for Mitra-v2 at your real scoring volume and put it next to the tree's CPU cost.
Before Next Budget Cycle:
- If a table beats your tree only with TabPFN-3.5, get a written Prior Labs quote for that table's volume. Until you have one, keep the tree in production.
Continue Reading
- Foundation Models Beat Tuned XGBoost Below 50,000 Rows
- NVIDIA's Kumo Tabular Undercuts TabPFN's License Below 60,000 Rows
- SAP's €1B Bet on the Other Half of Enterprise AI
- Nvidia's $400M Kumo Bet: LLMs Can't Touch Your Database
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
- Hugging Face Hired Bankers. Go Mirror Your Weights.
