TimesFM vs Chronos vs Moirai: The Leader's Weights Ban Production

Chronos-2 is the forecasting foundation model to self-host: Apache 2.0, P1 to P99 quantiles and a Decathlon deployment behind it. TimesFM 3.0 tops Google's benchmarks, but its weights are licensed for production only inside BigQuery, and Moirai 2.0's are non-commercial.

By Rajesh Beri·October 6, 2026·12 min read
Share:
A warehouse aisle of stocked pallet racks with a planner's laptop on a rolling cart in the foreground, the screen showing a weekly sales line chart with a shaded forecast band extending to the right.

Illustration generated using AI

Use Chronos-2 if you want to run a forecasting foundation model yourself, and use TimesFM through BigQuery's AI.FORECAST if your sales history already lives in BigQuery. Do not build on Moirai 2.0: its weights are licensed for non-commercial use only, it trails both rivals on the main public benchmark, and no major cloud sells it as a managed service. TimesFM 3.0 is the accuracy leader by Google's own launch numbers, but its downloadable weights ship under a licence that rules out production use, so for most buyers the model at the top of the leaderboard is one they can only rent.

The workload behind every comparison below: 25,000 weekly SKU-by-location series, three years (156 weeks) of history, a 52-week horizon, refreshed weekly, with P90 and P95 quantiles feeding a safety-stock calculation. That is a mid-sized retailer or distributor. Prices were checked on each vendor's live page on October 6, 2026.

Chronos-2 (Amazon) TimesFM 3.0 / 2.5 (Google) Moirai 2.0 (Salesforce) AutoETS / AutoARIMA (statsforecast)
Weights licence Apache 2.0 3.0: non-commercial, non-production. 2.5: Apache 2.0 CC BY-NC 4.0 Apache 2.0 (code)
Commercial production use Yes, anywhere 3.0 only via BigQuery; 2.5 anywhere No Yes
Max context / horizon 8,192 / 1,024 steps BigQuery 3.0: 2,048 / 1,024; BigQuery 2.5: 15,360 / 10,000 Not stated on model card Whole history
Quantiles 21, from P1 to P99 9, from P10 to P90 (BigQuery adds a 95% interval) Quantile loss, levels not stated Prediction intervals at any level
Covariates Past and known-future 2.5 and 3.0 via XReg; in BigQuery, 3.0 only (Preview) Not documented on card Model-dependent
Managed service SageMaker JumpStart, AutoGluon-Cloud BigQuery ML, Vertex Model Garden, Connected Sheets None None needed
Our workload, compute per weekly run About $0.05 on one CPU box Under $0.001 on-demand bytes (until Dec 1, 2026) Not licensable Cents
Verdict Default pick Pick if your data is in BigQuery Loser Keep as the baseline you must beat

What Is a Time Series Foundation Model?

A time series foundation model is a transformer pretrained on billions of time points from many domains, so it can forecast a series it has never seen without being trained on it ("zero-shot"). You pass it history; it returns a forecast and quantiles in one forward pass. The pitch to a demand planner is one model across every SKU, instead of one tuned ARIMA or Prophet model per series.

The three contenders here are the ones procurement teams ask about. TimesFM 3.0, published by Google Research on August 31, 2026, is a 330M-parameter decoder-only model pretrained on more than a trillion time points. Chronos-2, released by Amazon on October 20, 2025, is a 120M-parameter encoder model. Moirai 2.0 is Salesforce's, and its small variant has 11.4M parameters.

Which Model Is Most Accurate Zero-Shot?

On public benchmarks, TimesFM 3.0 leads, Chronos-2 is close behind, and Moirai 2.0 sits a tier lower. Treat every one of those rankings as a vendor claim until you have run your own backtest.

Google says TimesFM-3 has the best average rank among pretrained models on GIFT-Eval, FEV-Bench and TIME, ahead of Chronos-2 and Datadog's Toto 2.0, for both point and probabilistic forecasts. An independent write-up notes those figures come from launch material, not a peer-reviewed paper with task-level tables. Before TimesFM-3 shipped, Amazon's own paper reported that Chronos-2 beat TiRex and TimesFM-2.5 on GIFT-Eval in win rate and skill score. Salesforce's paper puts Moirai 2.0 5th on MASE and 6th on CRPS on the same benchmark, a large gain over Moirai 1.0 but still behind the two leaders.

Two caveats should shape how much weight you give the leaderboard. GIFT-Eval added a test-data leakage column in August 2025 because some models had been pretrained on data that overlapped the test split, so check the flag before you compare rows. And the benchmarks are dominated by series that look nothing like a sparse, promotion-driven SKU.

The production evidence that matters most for this buyer comes from Decathlon. In an AWS write-up dated August 28, 2026, the retailer backtested foundation models across 39,000 product series and 101 rolling cutoffs. Zero-shot Chronos-2 matched or beat its weekly-retrained production system. Fine-tuned Chronos-2 cut 12-week WAPE from 39% to 28% in Southeast Asia and from 53% to 38% in Latin America, and Decathlon says each point of WAPE is worth about 0.3 days of inventory.

Does Fine-Tuning Beat Zero-Shot?

Yes, by enough to matter, and only Chronos-2 has a public enterprise case showing it. Decathlon fine-tunes with LoRA once every six months and still gets several points of WAPE over zero-shot, so the refresh cadence is light.

Moirai's uni2ts library supports fine-tuning as well, but the licence makes that moot for a commercial deployment. BigQuery's AI.FORECAST is zero-shot only: you get Google's weights as they are, with no tuning step. If your SKUs behave unlike public data (intermittent spare parts, heavy promotions), that is the main thing you give up by renting TimesFM instead of running Chronos.

No model fixes a regime change. An observability team that tested Chronos, TimesFM, Toto and IBM's Tiny Time-Mixers on live metrics found that no model handled a first-of-its-kind event well zero-shot, though the pretrained models recovered faster than classical ones once the new pattern appeared. A new product launch or a pandemic-style demand shock still needs a planner's override.


Can You Legally Run Each One in Production?

Chronos-2 yes, Moirai 2.0 no, and TimesFM depends on the version and on where it runs. For most buyers this settles the purchase before accuracy does.

TimesFM's source code is Apache 2.0, but the TimesFM 3.0 weights are distributed under timesfm-non-commercial-license-v1.0, and the Hugging Face checkpoint carries the same licence. Google's BigQuery documentation states that using TimesFM through BigQuery is governed by the Google Cloud Terms of Service and allows commercial uses. The older TimesFM 2.5 weights are Apache 2.0, so you can self-host 2.5 commercially, just not the model Google's benchmarks are about.

Chronos-2 is Apache 2.0: download it, fine-tune it, run it on-prem or in another cloud, ship it inside your product. Moirai 2.0's weights are CC BY-NC 4.0, which bars commercial use. A demand plan that drives purchase orders is commercial use.

What Do Context Length and Horizon Limits Mean for You?

For weekly demand planning, none of the limits bind; they start to matter for daily or hourly data. Our workload has 156 points of history and a 52-step horizon, well inside every model.

Chronos-2 takes up to 8,192 steps of context and forecasts up to 1,024. In BigQuery, the AI.FORECAST reference caps TimesFM 3.0 at a 2,048-step context and a 1,024-step horizon, while TimesFM 2.5 accepts context windows up to 15,360 and horizons up to 10,000. So the newer model sees less history in BigQuery. Five years of daily sales (about 1,825 points) fits under 3.0's cap; a year of hourly energy load (8,760) does not, and for that workload 2.5 or Chronos-2 is the better fit.

Do You Get the Quantiles a Safety-Stock Formula Needs?

Chronos-2 gives you the tails directly; raw TimesFM stops at P90. Chronos-2 predicts 21 quantiles from 0.01 to 0.99, including P95 and P99. TimesFM returns nine quantiles from 0.1 to 0.9. If your safety stock targets a 95% service level and you self-host TimesFM, you are extrapolating past the model's top quantile.

BigQuery papers over this: AI.FORECAST takes a confidence_level that defaults to 0.95 and returns prediction-interval bounds. Check how those bounds behave on your own intermittent SKUs before you wire them into reorder points.


Where Does Each One Run?

TimesFM is the easiest to adopt if you are a Google Cloud shop, Chronos-2 the most portable, and Moirai runs only where you put it.

TimesFM is built into BigQuery ML in every BigQuery region through a single SQL function, available in Vertex Model Garden, and since July 1, 2026 from Connected Sheets. For a planning team that lives in SQL and spreadsheets, there is nothing to deploy. Model Garden is the route if you want a dedicated Google Vertex AI endpoint instead.

Chronos-2 has been on SageMaker JumpStart since December 30, 2025 and AutoGluon-Cloud since June 5, 2026. JumpStart deploys it as a real-time endpoint on CPU or GPU instances, which a weekly batch does not need. Decathlon skips the endpoint and runs inference on a plain m6i.8xlarge CPU instance, using a g5.4xlarge only for the twice-yearly fine-tune. The weights are on Hugging Face, so the same code runs on Azure, on Databricks or on-prem.

Moirai 2.0 is a Hugging Face download plus Salesforce's research library. No hyperscaler sells it as a managed service.

What Does a Forecast Cost Against Your ARIMA or Prophet Pipeline?

Compute is a rounding error for all of them; the money is in the people who tune per-series models. At our workload, every option costs cents per weekly run.

For Chronos-2, Decathlon reports inference on 15,000 series in about 75 seconds and roughly $0.03 per weekly run. Scaling linearly to our 25,000 series gives about 125 seconds on an m6i.8xlarge, which lists at $1.536 an hour on-demand in us-east-1. That is about $0.05 a run, or under $3 a year before the fine-tuning job.

For TimesFM in BigQuery, on-demand queries bill by bytes processed at $6.25 per TiB after the first free TiB each month. Our 3.9 million rows of history come to well under a gigabyte, so the monthly forecast runs fit inside the free tier on a project that does nothing else. On Enterprise editions it draws on your slot commitment instead. That changes for TimesFM 3.0: the AI.FORECAST reference says it moves to token-based pricing from December 1, 2026. Read the token rate on the pricing page before you budget a 3.0 rollout past November.

The classical stack is just as cheap to run. Nixtla's statsforecast (Apache 2.0) claims it fits a million series in 30 minutes on Ray and runs 500x faster than Prophet, both vendor figures. Prophet itself has been in maintenance mode since v1.4.0, taking only bug fixes. If Prophet is your baseline, you are replacing it whichever way you go.

The cost that does move is labour. Decathlon says opening a new region dropped from about six months to two or three once one pretrained model replaced its per-region pipeline. That and the WAPE gain make the business case.


Who Should Not Pick Each Option?

Each option has a buyer it fails.

Skip Chronos-2 if you have no one to own a Python batch job. It is the most flexible option and also the one that needs an engineer, an instance and a scheduler. A planning team with only SQL access will get further, faster, with BigQuery.

Skip TimesFM 3.0 if your data is not in BigQuery or you need to run on-prem, in another cloud or inside a product you sell. The downloadable weights do not cover you, and copying your sales history into BigQuery just to rent a forecast is a data-platform decision that a forecasting tool should not drive. Skip it also if you need a fine-tuned model; BigQuery gives you none.

Skip Moirai 2.0 for any commercial forecast. It is the loser here: a non-commercial licence, a 5th/6th place on GIFT-Eval behind both rivals, and no managed service. It is a sound research model and a fine thing to benchmark against. Nothing about it justifies a production build.

Skip a foundation model entirely if you forecast a few hundred stable series that your current AutoETS already gets right, or if you need sub-second, always-warm inference. The Parseable team found classical methods still preferable for stable, stationary workloads. A 2025 study of Chronos against standard methods found Chronos's edge grew with horizon length, which also means a short-horizon, well-tuned statistical model leaves less to gain.

What Should Decide Your Choice?

Three questions predict regret better than any benchmark: where your data already sits, whether you will ship the model outside your own walls, and whether a fine-tune is worth an engineer.

  • If your data is in BigQuery and your team writes SQL, use AI.FORECAST with TimesFM 2.5 now, test 3.0 with covariates, and re-check the price before December 1.
  • If your data is in S3, Snowflake, Databricks or on-prem, run Chronos-2 on a CPU instance as a weekly batch and plan a LoRA fine-tune within six months.
  • If the forecast goes into a product you sell to customers, use Chronos-2, or TimesFM 2.5 if you want Google's lineage. Neither TimesFM 3.0 nor Moirai 2.0 is licensed for it.
  • If you forecast hourly data with long seasonal cycles, use TimesFM 2.5 in BigQuery (15,360 context) or Chronos-2 (8,192). TimesFM 3.0's 2,048 cap in BigQuery is too short.

What changes the answer: Google releasing TimesFM 3.0 weights under Apache 2.0 (as it did for 2.5) would make it the default self-host pick. A Chronos-3 that closes the benchmark gap would settle it the other way.

What to Do Next

This Week: Pull 18 months of history for your 500 highest-revenue SKUs and score your current production forecast on WAPE and on P90 coverage. You need that number to judge any vendor's result.

This Month: Backtest Chronos-2 zero-shot on the same SKUs with at least 12 rolling cutoffs, and if you are on Google Cloud run AI.FORECAST alongside it. Ask your legal team to read the TimesFM 3.0 licence before anyone downloads the checkpoint into a shared environment.

Before Q1 Close: If either model beats your baseline by more than two WAPE points, move one region onto it in parallel with the old system for a full quarter, keep planner overrides for launches and promotions, and schedule the first fine-tune.

The Bottom Line

The pattern matches what happened to tabular data a year earlier: a pretrained model now beats tuned trees on small tables, and the licence on the best checkpoint decides who can actually use it, as with Mitra and TabPFN. For forecasting, Chronos-2 is the open model with production evidence behind it, TimesFM is the managed option for BigQuery customers, and Moirai is for research. Run the backtest on your own SKUs before you sign anything that commits a region.

Continue Reading

Share:

Frequently Asked Questions

Is TimesFM 3.0 free for commercial use?

Not as a download. Google distributes the TimesFM 3.0 weights under timesfm-non-commercial-license-v1.0, which rules out commercial and production use. Using TimesFM through BigQuery's AI.FORECAST function is governed by the Google Cloud Terms of Service and allows commercial use. The older TimesFM 2.5 weights are Apache 2.0.

Which time series foundation model is best for demand forecasting?

For most teams, Chronos-2. It is Apache 2.0, supports past and known-future covariates, outputs 21 quantiles from P1 to P99, and Decathlon reports that a fine-tuned Chronos-2 cut 12-week WAPE by 11 to 15 points against its legacy system. If your data already lives in BigQuery, TimesFM via AI.FORECAST is the simpler option.

Can I use Salesforce Moirai 2.0 in production?

Not for commercial work. The Moirai 2.0 weights on Hugging Face are licensed CC BY-NC 4.0, which bars commercial use, and no major cloud offers it as a managed service. It ranked 5th on MASE and 6th on CRPS on GIFT-Eval in Salesforce's own paper, behind Chronos-2 and TimesFM.

How much does a forecasting foundation model cost per forecast compared with ARIMA or Prophet?

Compute costs cents for all of them. Decathlon runs Chronos-2 inference on a CPU instance for about $0.03 per weekly run, and BigQuery bills TimesFM on-demand by bytes processed at $6.25 per TiB after a free monthly TiB. The real saving is labour: Decathlon cut new-region rollout from about six months to two or three.

What are the context and horizon limits of TimesFM and Chronos-2?

Chronos-2 accepts up to 8,192 steps of context and forecasts up to 1,024 steps. In BigQuery, TimesFM 3.0 is capped at a 2,048-step context and a 1,024-step horizon, while TimesFM 2.5 accepts up to 15,360 steps of context and a 10,000-step horizon.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →