NVIDIA's Kumo Tabular Undercuts TabPFN's License Below 60,000 Rows

NVIDIA's Kumo Tabular ships under OpenMDW-1.1 for commercial use, while TabPFN-3.5's open weights stay non-commercial. Its #1 TabArena Elo is NVIDIA's own run, and it was pretrained only to 60,000 rows and 100 columns.

By Rajesh Beri·September 29, 2026·8 min read
Share:
A data scientist's desk at night with a single workstation GPU card lying beside a thick printed spreadsheet whose rows run off the edge of the desk, a signed paper licence agreement weighted down by the GPU.

Illustration generated using AI

If your churn, fraud or credit-risk model uses under 60,000 rows and 100 numeric or categorical columns, you can now replace tuned XGBoost with a tabular foundation model without a commercial licence. NVIDIA's Kumo Tabular, published on 29 September, ships its weights under OpenMDW-1.1 "for commercial use" — while Prior Labs' TabPFN-3.5, released two weeks earlier, keeps its open weights non-commercial. What you should not do is treat NVIDIA's #1 leaderboard claim as settled: it comes from NVIDIA's own evaluation run, and TabArena's maintainers have not finished reproducing it.

That leaves a licensing decision you can make this quarter and an accuracy question you should answer on your own data.

What NVIDIA Actually Released

Kumo Tabular is a pretrained transformer that reads a labelled table as context and predicts the labels of new rows in a single forward pass — no training, no hyperparameter search. NVIDIA's launch post describes three sizes spanning 28M to 215M parameters, pretrained entirely on synthetic tables generated from structural causal models.

A tabular foundation model is a model pretrained once on millions of tables so that it can make predictions on a new table by in-context learning, instead of fitting a fresh model per dataset.

The code lives in NVIDIA's structured-data-models repository under Apache 2.0, with a fit/predict interface that will look familiar to anyone who has used scikit-learn, and a GPU is expected — the repo "highly recommend[s]" cuDF. The weights sit on the Kumo-Tabular model card under OpenMDW 1.1. If you already pull models through Hugging Face, there is nothing new to procure.

The name is not decoration. NVIDIA acquired Kumo AI in June, bringing over co-founders Vanja Josifovski, Hema Raghavan and Jure Leskovec, whose relational foundation model made churn and credit-default predictions without per-task training. We covered what that deal meant for relational data four months ago; Kumo Tabular is the single-table model now shipping under that name.

Why the Licence Is the Real News

The licence is the part that changes a procurement decision, because the accuracy gap between the top models is small and the licence gap is binary.

OpenMDW-1.1 grants royalty-free permission to "deal in the Model Materials without restriction," covering copyright, patent, database and trade-secret rights. It explicitly imposes no restrictions on outputs. The conditions are light: keep the licence and origin notices with any redistribution, and lose your rights if you sue claiming the model materials infringe your IP. The Linux Foundation released 1.1 on 28 May, with NVIDIA adopting it across Cosmos, Isaac GR00T, Ising and Nemotron — so your legal team may already have reviewed it once.

Compare the incumbent. The TabPFN-3.5 weights carry tabpfn-3-5-license-v1.0, which allows research and internal evaluation but states that "the model, its derivatives, and its outputs cannot be used for any commercial or production purpose." Production means a Commercial Enterprise License via Prior Labs sales. The pricing page lists Free, Pro and Max API tiers plus commercial Private Cloud and on-prem licences — and no published rates for any of them.

Read the predecessor's terms too, because they show how far "non-commercial" reaches. The TabPFN-2.5 licence listed revenue-generating products, client deliverables, internal commercial decision-making and competitive benchmarking for procurement as out of bounds. If your data science team ran a TabPFN bake-off against XGBoost on production data to justify a purchase, your counsel should check whether the 3.5 terms say the same thing.

Prior Labs is also no longer an independent startup. SAP completed its acquisition on 17 July, committing more than €1 billion over four years while Prior Labs operates as an independent entity, and on 15 September put TabPFN-3.5 Plus into SAP AI Core. The context is in our piece on the SAP deal. So the choice in front of an ML platform team is now NVIDIA's permissive weights against SAP's paid model — two platform vendors, each with a reason to pull your tables toward its stack.


Both Vendors Claim First Place on TabArena

Both NVIDIA and Prior Labs say they rank first on TabArena, and neither headline number has been produced by the benchmark's maintainers.

NVIDIA reports Kumo Tabular at an Elo of 1950 on TabArena, plus first place on BeyondArena (Elo 1418), TALENT and ScoringBench, running 17 times faster than LimiX-2. The comparison was "all three Kumo Tabular sizes with default settings against the full TabArena leaderboard," on "a uniform single RTX 6000 Pro evaluation setup."

Prior Labs reports TabPFN-3.5-Thinking at 1910 Elo and base TabPFN-3.5 at 1866. TabArena's maintainers added base TabPFN-3.5 as a verified model on 22 September; the Thinking variant behind the 1910 is not on that verified list. Prior Labs' own release notes still describe the family as "ranking first on TabArena and BeyondArena" and first on ScoringBench.

Here is why that matters. TabArena's design is that a team of maintainers runs the models itself, and its paper shows validation method and ensembling change results. The pull request adding Kumo Tabular to the benchmark, opened on 28 September by TabArena co-author Lennart Purucker, was still open when we checked, reading "Maintainer runs in progress" with sign-off "to be requested from NVIDIA." Until that lands, 1950 is NVIDIA's number, measured NVIDIA's way.

A 40-point Elo gap between two vendor-reported numbers is not a result you should buy on. We made the same argument about agent benchmarks with small gaps: most differences that size disappear once you look at the confidence interval.

Where Kumo Tabular Stops

The constraints decide who can use it more than the leaderboard does.

NVIDIA's own post sets four limits:

  • Rows and columns. The final pretraining stage "extends it to 60,000 rows, still with up to 100 columns." NVIDIA warns that accuracy "may degrade on tables far beyond the training ranges."
  • Classes. "A single forward pass covers up to 10 classes," which the library stretches to more classes with error-correcting output codes — a workaround, not native support.
  • Data types. Numerical and categorical columns only. Text, images and timestamps have to be turned into features through preprocessing recipes first.
  • Distribution shift. Accuracy may also degrade "when the query rows come from a different distribution" than the context rows — which is the normal condition for a fraud model six months after deployment.

Now set that against TabPFN-3.5's stated envelope: up to 1,000,000 rows (subject to feature count and API limits) and up to 20,000 features, with 6,000 recommended. Its technical report claims handling of strings, text, images, high-cardinality categoricals and non-i.i.d. temporal splits. On breadth, the paid model is still well ahead.

Steel-man NVIDIA's position: many enterprise models really are small. Risk scorecards, per-segment churn models and claims triage often sit well under 60,000 rows, and our September explainer on the crossover found that is exactly where foundation models beat tuned trees. For that slice, a free commercial licence beats a better envelope you have to pay for.

What the Leaderboard Does Not Price

A foundation model is free to fit and costs money to serve, and neither release publishes the number you need.

NVIDIA's evaluation ran on a single RTX 6000 Pro; the TabArena maintainer runs use RTX PRO 6000 (96 GB) workers per the pull request. XGBoost runs on CPU. Every prediction from Kumo Tabular re-reads its context table on a GPU, so your cost per thousand predictions depends on context size, batch size and whether you reuse cached context — the repo supports cached inference workflows, and Prior Labs charges 75% less for successful KV-cache reuse on its API. Neither vendor gives you a cost per prediction at your volume. You will have to measure it.

The other unpriced item is model drift on the vendor side. Open weights mean you can pin a version and mirror it, which is the discipline we laid out after Hugging Face's ownership questions. An API model can change underneath your validated scores.

What to Do Before You Sign a Tabular Licence

This Week:

  1. Inventory your tabular models by shape. For each production model, record rows, columns, class count and whether it uses free text or timestamps. Only models under 60,000 rows, 100 columns and 10 classes are in Kumo Tabular's native envelope.
  2. Send the TabPFN-3.5 licence and OpenMDW-1.1 to counsel together. Ask one question: is our current TabPFN evaluation on production data inside "internal evaluation," or is it "commercial decision-making"?

This Month:

  1. Run a three-way bake-off on two in-envelope models. Kumo Tabular (default settings), your tuned XGBoost or CatBoost, and TabPFN-3.5 if your licence allows it. Use a time-based split, not a random one — the drift warning is the one that bites in production.
  2. Measure GPU cost per 1,000 predictions at your real scoring volume, with and without cached context, and put it next to the CPU cost of the tree you already run.
  3. Mirror and pin the weights into your internal model registry with a hash, so the validated version is the one you deploy.

Before Renewal:

  1. Watch TabArena pull request #625. If the maintainers' run confirms Kumo Tabular near the top, use it as leverage in any Prior Labs or SAP AI Core negotiation. If it does not, you have lost nothing by waiting.

The Bottom Line

This is the pattern open-weight language models set in 2024 and 2025: the proprietary leader keeps the breadth, an open challenger gets close enough on the common case, and the licence decides most deployments. On tabular data the common case is small tables, and for those NVIDIA just removed the licence fee. The breadth — a million rows, text columns, wide tables — still belongs to TabPFN-3.5, and to SAP.

The leaderboard will settle itself within weeks. The licence question is yours to settle now.

Continue Reading

Share:

Frequently Asked Questions

Can Kumo Tabular be used commercially?

Yes. NVIDIA released the Kumo Tabular weights under the OpenMDW-1.1 licence for commercial use. It grants royalty-free rights without restriction, imposes nothing on outputs, and requires you to keep the licence and origin notices when redistributing; rights terminate if you sue claiming the model infringes your IP.

Is TabPFN-3.5 free for commercial use?

No. The open TabPFN-3.5 weights carry tabpfn-3-5-license-v1.0, which allows research and internal evaluation but bars any commercial or production use of the model, its derivatives and its outputs. Production requires a Prior Labs commercial licence or its paid API.

What are Kumo Tabular's limits?

NVIDIA says its final pretraining stage covers tables up to 60,000 rows and 100 columns, a single forward pass handles up to 10 classes, and it supports only numerical and categorical columns. Accuracy may degrade on tables far beyond those ranges or when query rows come from a different distribution.

Is Kumo Tabular really first on TabArena?

NVIDIA reports an Elo of 1950 from its own run on a single RTX 6000 Pro, against a self-reported 1910 for TabPFN-3.5-Thinking. As of 30 September 2026 the TabArena pull request adding Kumo Tabular was still open with maintainer runs in progress, so the ranking is not yet independently reproduced.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →