KumoRFM vs XGBoost: Its 14-Point Win Beat Someone Else's Model

KumoRFM's 14-point RelBench lead was over a LightGBM that saw only one table. Keep XGBoost in production, test KumoRFM against your own model, and get the NIM licence in writing first.

By Rajesh Beri·October 8, 2026·14 min read
Share:
A data analyst's desk at night with a large printed database schema diagram of a dozen linked tables spread across it, connected by drawn lines, next to a laptop showing a single scoring table.

Illustration generated using AI

If you are about to fund another six-month feature-engineering project, stop and run KumoRFM against your current production model first. If you are about to retire that model because of a zero-shot benchmark, don't do that either. KumoRFM's headline result, 76.71 average AUROC zero-shot against 62.44 for LightGBM on RelBench, compares it to a baseline built for a paper. Nobody has published how it does on a warehouse like yours. The production container also carries a licence that, as of October 9, 2026, restricts the model to non-commercial research or evaluation.

Our verdict: keep XGBoost (or whatever gradient-boosted model you run today) as the production default. Use KumoRFM as a fast challenger on the long tail of prediction tasks you never staffed. Promote it only where it beats your own model on a temporal holdout and legal has signed off on the licence you are actually deploying.

Licences, prices and availability below were checked on each vendor's live page on October 9, 2026.

Option Reads many tables natively Production licence (Oct 9, 2026) Published price Time to first prediction Verdict
NVIDIA Kumo Relational (KumoRFM) Yes, via the join graph NGC NIM: non-commercial research/eval only; Hugging Face weights: OpenMDW 1.1 None published; contact NVIDIA About 1 second per query (vendor claim) Best challenger; not yet a replacement
XGBoost + your feature pipeline No, you flatten Apache-2.0 Free; you pay in engineers About 12.3 hours and 878 lines per task (vendor-run study) Keep as the production default
AutoGluon No, you flatten Apache-2.0 Free Your training time budget, after features exist Use it to raise your baseline before any bake-off
Open-source RDL (RelBench + PyTorch Geometric) Yes MIT Free One GNN training run per task Only if you employ GNN engineers
TabPFN-3.5 No, single table Open weights non-commercial; commercial is contact sales None published Seconds, after flattening Loser for multi-table work

What Does a Relational Foundation Model Do That XGBoost Cannot?

A relational foundation model predicts directly from several linked tables, using the primary and foreign keys as structure, instead of requiring you to flatten them into one feature table first. KumoRFM treats each row as a node and each key relationship as an edge, samples historical entities with known labels from your own database, and uses those as in-context examples for the entities you want scored. No per-task training happens in the zero-shot mode.

XGBoost needs one row per entity. To predict churn for a customer, someone writes SQL that rolls orders, line items, tickets and payments up into "orders in the last 30 days", "mean ticket resolution time" and a few hundred siblings. That flattening is where the work and the risk sit. It discards cross-table structure you didn't think to aggregate, and it's where label leakage gets in: a feature computed with data from after the prediction date.

A tabular foundation model such as TabPFN sits in between. It replaces the XGBoost training step, but it still expects the single flattened table, so it saves you nothing on the pipeline that costs the most. (Prior Labs' TabPFN-REL adds automatic flattening and, in Prior Labs' own report, edges KumoRFM-2 on RelBench, but that is a research add-on, not the product you license.) That is why it loses this comparison even though it is strong on small single tables (we covered that fight in our Mitra vs TabPFN vs XGBoost comparison).

What Do the Published Numbers Actually Show?

The published numbers show KumoRFM matching or slightly beating supervised relational models on public benchmarks built partly by Kumo's own researchers, and beating gradient boosting on raw single-table features by a wide margin. Read them with that authorship in mind.

On RelBench, which spans seven databases, 30 tasks, 51 tables and 103,466,370 rows, the first KumoRFM paper reports average AUROC across 12 entity-classification tasks of:

  • 76.71 for KumoRFM zero-shot, which never trained on any RelBench dataset
  • 75.83 for a supervised graph neural network trained per task
  • 62.44 for LightGBM on the entity table's raw features alone
  • 81.14 for KumoRFM after fine-tuning

The KumoRFM-2 paper raises the zero-shot average to 79.60, against 78.06 for RelGNN, the strongest supervised relational model, and 63.67 for LightGBM. Note that the LightGBM figure is not the same in the two papers. On regression the picture is weaker: the first paper reports the data scientist baseline leading three of nine regression tasks.

On SAP's SALT benchmark, built from a real ERP system, KumoRFM-2 reports a mean reciprocal rank of 0.83 in-context and 0.89 fine-tuned, against 0.77 for a data scientist's features fed to AutoGluon and 0.79 for CARTE. Kumo's April 14, 2026 launch release repeats the fine-tuned and baseline figures, but it also gives two different margins over supervised models (5% and 1.5%) in the same document, and claims scale to "500 billion+ rows" without describing a test.

Three things about provenance matter for a buyer:

  1. RelBench came from the same group. The RelBench paper is by researchers at Stanford and Kumo.AI, and SALT was folded into RelBench v2 in January 2026, per the RelBench repository. That does not make the results wrong, and the benchmark is public and MIT-licensed. It does mean every headline number was produced by the vendor on a benchmark its own researchers designed.
  2. The only outside re-run comes from a competitor. Prior Labs, which sells TabPFN, re-ran KumoRFM-2 with the authors' scripts in its TabPFN-3 technical report (May 2026). In that run KumoRFM-2 trails the supervised RelGNN on classification, and the report says KumoRFM's first version likely followed an evaluation regime that overestimates performance, especially on the rel-f1 tasks. A later RelArena-α paper raises the same rel-f1 concern. We found no neutral replication.
  3. The 14-point gap is against a weak baseline. LightGBM at 62.44 sees only the entity table's own columns, nothing from related tables. The fairer comparison is the supervised GNN, and there the zero-shot lead is under one point in the first paper and about 1.5 points in the second.

Tree models have a long record of beating deep learning on ordinary tabular data, documented in Grinsztajn, Oyallon and Varoquaux's study. KumoRFM's claim is narrower: it wins once the signal lives across tables. That claim is plausible, but the only outside test so far comes from a rival vendor, and that test puts KumoRFM-2 behind the best supervised model.

How Much Faster Is It, Really?

On time to first prediction, KumoRFM wins by orders of magnitude, and this is the advantage most likely to survive your data. The KumoRFM paper puts a zero-shot answer at about one second and one line of Predictive Query Language, against about 30 minutes and 56 lines for a supervised GNN pipeline and about 12.3 hours and 878 lines of code for the manual data scientist workflow.

Those are vendor-measured, and the 12.3 hours understates what your team actually spends: it excludes the weeks of label definition, stakeholder argument and monitoring setup that come before and after. The RelBench paper, also co-authored by Kumo researchers, found relational deep learning cut human hours by 96% and lines of code by 94% against manual feature engineering, which points the same way.

To make the comparison concrete, picture this workload: a 12-table commerce warehouse (customers, accounts, orders, line items, products, returns, payments, invoices, tickets, sessions, campaigns, contracts), about 50 million rows, three tasks (90-day churn, next-30-day spend, a product recommendation), scored weekly in batch.

  • XGBoost: one feature pipeline per task, built and maintained by someone who knows the schema, plus drift monitoring for each.
  • KumoRFM: three PQL statements. A query looks like PREDICT SUM(orders.price, 0, 30, days) FOR EACH users.user_id, per the client's README. The hard part becomes the schema metadata: correct keys, timestamps and enough history to fill the window.
  • AutoGluon: same flattening work as XGBoost, then a better-tuned ensemble on top. The pipeline work stays with you.

Where speed changes the economics is the long tail. Most enterprises have dozens of prediction questions (which suppliers will ship late, which invoices will be disputed, which accounts will downgrade) that never cleared the bar for a dedicated pipeline. For those, a one-second answer with roughly supervised-level accuracy beats no model at all.


Where Does Your Data Go?

Where your data goes depends on which of three Kumo surfaces you use, and only one of them keeps it fully in your warehouse by design.

  1. Snowflake Native App (Kumo Predict). NVIDIA's documentation says the app runs on Snowpark Container Services inside your Snowflake account, and that Kumo the company never has direct access to training data, model artifacts or prediction tables. You pay Snowflake for compute plus whatever the provider charges, under Snowflake's two-layer native app cost model. No Kumo rate card is published.
  2. NIM endpoint (KumoRFM zero-shot). The open-source kumo-relational-client (Apache-2.0) ships connectors for Snowflake, Databricks, PostgreSQL, S3, DuckDB and SQLite. It sends the rows to score to a NIM endpoint. Use NVIDIA's hosted endpoint and the rows leave your environment. Self-host the container on your own A10G, L4, L40S or H100 GPUs, per the NGC listing, and they don't. The optional [explain] extra posts row data to a third-party LLM endpoint, so leave it off until security has reviewed it.
  3. Hugging Face weights. nvidia/Kumo-Relational runs in your own Python process on a CUDA device. Nothing leaves, but you own the serving.

If you run Databricks rather than Snowflake, the client's Databricks connector is the route; we found no Databricks-native app equivalent to the Snowflake one. For the broader platform choice, see our Snowflake Cortex vs Databricks Mosaic AI comparison and the tool pages for Snowflake Cortex AI and Databricks Mosaic AI.

What Did NVIDIA's Acquisition Change?

NVIDIA's acquisition changed the licensing and the front door, and buyers should treat Kumo as an NVIDIA product line whose commercial terms are still settling. Fortune reported on June 3, 2026 that NVIDIA had acquired Kumo, at a price The Information reported as $400 million (we covered the deal in Nvidia's $400M Kumo Bet).

What we can verify today:

NVIDIA is shipping the product, so it is unlikely to disappear. The practical risk is building on a surface (the Snowflake app, the hosted endpoint) whose price and terms get set after you depend on it. Get a written quote and licence for the exact artifact before production, and keep the XGBoost pipeline warm until you have it. Our Kumo Tabular licence piece covers the same split on the single-table model.

Who Should NOT Pick Each Option?

Each option has a buyer it fails, and naming them is more useful than another feature row.

  • Skip KumoRFM if your target lives in one table already (TabPFN, Mitra or XGBoost will do), if your keys and timestamps are unreliable (the model's context is only as good as the join graph), if your task is a high-error regression where the paper itself concedes expert features can win, or if you need commercial production use this quarter and cannot get the NIM licence in writing.
  • Skip XGBoost as your only plan if you have more than a handful of untouched prediction questions and no headcount to build pipelines for them. At 878 lines per task, most of those questions will stay unanswered.
  • Skip AutoGluon if your bottleneck is the feature pipeline. It improves the model on top of features you still have to write. Use it to set an honest baseline, which is exactly how the SALT comparison used it.
  • Skip open-source RDL if nobody on your team has trained a graph neural network. The RelBench code is free and MIT-licensed, and you will spend the savings on GPU debugging.
  • Skip TabPFN for this job whatever your data size. The licensed model doesn't touch the multi-table problem, and production use outside the API needs a commercial licence from Prior Labs with no published price.

How Do You Run a Fair Bake-Off on Your Own Schema?

A fair bake-off compares KumoRFM to the model you run in production today, on a temporal split, with a leakage audit, on a label definition the business has signed. Anything else reproduces the paper's advantage.

  1. Pick three tasks: one you already model well, one you model badly, one you never modelled. The first tells you whether KumoRFM can match production; the third is where it most likely wins.
  2. Fix the label definition first and in writing. "Churn" as 90 days without an order and "churn" as a contract non-renewal are different problems. Label definition moves results more than model choice, and it is the variable no public benchmark controls for your business.
  3. Split by time, never at random. Train or provide context up to a cutoff date, score the following window, and repeat for at least three cutoffs. The KumoRFM-2 paper uses anchor-time replay for the same reason.
  4. Audit leakage on both sides. Check every XGBoost feature and every table KumoRFM can reach for columns written after the anchor time: status fields that get updated, "last_modified" stamps, backfilled values. Kapoor and Narayanan's survey found leakage behind inflated results in published ML work across many fields, and a relational model that can see every table can see more of it.
  5. Report the metric the business uses. AUROC hides class imbalance. For a 2% churn base rate, report precision at the top decile you can actually act on.
  6. Use your production model as the baseline, then AutoGluon on the same features. If KumoRFM can't beat the better of the two by a margin larger than your cutoff-to-cutoff variance, it hasn't won.

This Week: Ask NVIDIA for the commercial licence terms for the Kumo Relational NIM and for the Snowflake app, in writing. Pick the three tasks.

This Month: Self-host the NIM or install the Snowflake app in a non-production account. Run the zero-shot pass on all three tasks across three time cutoffs, alongside your production model.

Before Q1 Close: Promote KumoRFM on any task where it wins on every cutoff and the licence is signed. Keep XGBoost on the rest and drop the new pipeline project only for the tasks KumoRFM now covers.

The Bottom Line

The relational foundation model is a real change in how much it costs to get a first prediction out of a multi-table warehouse. It has not yet shown that the prediction beats the one you already ship. The evidence comes from vendors (Kumo, and one re-run by a rival), on benchmarks Kumo's researchers helped build, against baselines weaker than a mature production model.

AutoML tools such as DataRobot followed a similar path: quickest to a decent model, and used most on the tasks nobody had time to hand-build. Expect KumoRFM to cover the long tail first and win the headline tasks one bake-off at a time, if it wins them at all. Send the licence request to NVIDIA before you schedule the first test.

Continue Reading

Share:

Frequently Asked Questions

Is KumoRFM better than XGBoost?

On RelBench, zero-shot KumoRFM averaged 76.71 AUROC against 62.44 for LightGBM on single-table raw features, but under one point ahead of a supervised graph model. Kumo produced those results; the only outside re-run, by rival Prior Labs, put KumoRFM-2 behind the supervised RelGNN. Nobody has published results on a private warehouse, so test it against your own production model before replacing anything.

Can I use the Kumo Relational NIM commercially?

As of October 9, 2026, NVIDIA's NGC listing for the Kumo Relational NIM says the model is licensed for non-commercial research or evaluation only. The Hugging Face weights use OpenMDW 1.1. Get written commercial terms from NVIDIA for the exact artifact you deploy.

Does my data leave Snowflake when I use Kumo?

With the Snowflake Native App, NVIDIA's documentation says it runs on Snowpark Container Services in your account and Kumo never accesses your data. With the NIM client, the rows you score go to whichever endpoint you call, so self-host the container if data cannot leave.

How much does KumoRFM cost?

No public price exists. As of October 9, 2026, kumo.ai and its pricing page redirect to NVIDIA documentation, and the Snowflake app bills through Snowflake compute plus provider fees with no published Kumo rate. Treat it as contact sales.

How do I run a fair test of a relational foundation model?

Use time-based splits across at least three cutoffs, audit every table for columns written after the prediction date, fix the label definition in writing, report the business metric rather than AUROC alone, and compare against your current production model and AutoGluon on the same features.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →