If your task is annual and a 10 m pixel resolves it, start with Google's AlphaEarth Foundations embeddings, pin the dataset version, and train gradient-boosted trees on hundreds to thousands of your own labels. Fine-tune Prithvi-EO-2.0 only when you need pixel-level segmentation from 30 m time series that you control end to end. Treat Clay v1.5 as the specialist pick for mixed sensors, aerial imagery included. Keep paying a commercial imagery vendor only for the jobs that need daily revisit or sub-metre pixels, because none of these models can manufacture detail the input never captured.
The reader this is for: you run data for an insurer, a utility or an agribusiness, a per-scene imagery analytics contract is up for renewal, and the question is whether your own labels plus an open embedding layer can do the same work for a fraction of the cost. For most annual scoring jobs they can. The catch is what happens when the embedding space changes underneath a model you already shipped.
| AlphaEarth Foundations | Prithvi-EO-2.0 | Clay v1.5 | Commercial per-scene analytics (Planet Insights Platform) | |
|---|---|---|---|---|
| What you get | Precomputed 64-dimension embeddings, one layer per year | Model weights (300M or 600M, plus TL variants) | Model weights (632M) and published embeddings | Imagery plus vendor-derived analytics |
| Input resolution | 10 m pixels | 30 m HLS | Sentinel-2, Landsat, Sentinel-1, NAIP, LINZ, MODIS | 3 m PlanetScope, 50 cm SkySat |
| Temporal grain | Annual | Multi-temporal time series | Per-chip, with time and location inputs | Daily or near-daily |
| Licence | CC-BY 4.0 | Apache 2.0 | Apache 2.0 (code and weights) | Vendor EULA; derived products governed by contract |
| Model cost | $0 | $0 (you pay for GPUs) | $0 (you pay for GPUs) | No public per-scene rate |
| Who should NOT pick it | Anyone who needs sub-annual or sub-10 m answers | Teams with no GPU or ML engineering capacity | Teams that need one steward promising stability | Anyone whose task is annual and 10 m |
| Verdict | Default for annual risk scoring | Pick for segmentation you will maintain | Specialist, multi-sensor only | Keep only for daily or sub-metre work |
Prices and terms were checked on each vendor's live page on October 11, 2026.
What Each Option Actually Delivers
A geospatial foundation model is a network pretrained on large volumes of unlabelled satellite imagery so that its outputs, called embeddings, already encode surface properties; you then fit a small model on your labels instead of training a vision pipeline from scratch. The three open options differ mainly in who runs the network.
AlphaEarth Foundations runs it for you. Google publishes the Satellite Embedding V1 dataset as annual layers of 10 m pixels, each pixel carrying 64 bands, one per embedding axis, in Earth Engine and in the Cloud Storage bucket gs://alphaearth_foundations. The same layers are mirrored as Cloud-Optimized GeoTIFFs on Source Cooperative, which lists coverage from 2017 to 2025 and notes that "This copy is not officially supported by Google." The AWS Registry of Open Data entry for that mirror carries the same CC-BY 4.0 licence. You never touch a GPU.
Prithvi-EO-2.0 hands you the weights. IBM and NASA publish them on Hugging Face under Apache 2.0, pretrained on 4.2 million samples of NASA's Harmonized Landsat Sentinel-2 product at 30 m, in variants from a 5M-parameter "tiny" to 600M, with "TL" versions that take time and location as inputs. You fine-tune and serve it yourself. Our Hugging Face tool page covers the hosting side.
Clay v1.5 also hands you weights, plus some published embeddings. Its model specification lists 632M parameters, training on 70 million chips from Sentinel-2, Landsat, Sentinel-1 SAR, LINZ, NAIP and MODIS, and a weights release dated 2024/11/19, all Apache 2.0. The Hugging Face model card says Clay "is an initiative of Renaissance Philanthropy" and that its embeddings are kept on Source Cooperative under ODC-BY.
The commercial incumbent sells something different: fresher, sharper pixels. NASA's vendor summary for Planet describes PlanetScope imaging the land surface near-daily at 3 m and SkySat at 50 cm. If your contract buys you that, no open model replaces it.
Do Embeddings Plus Your Labels Match a Bespoke Pipeline?
Yes, for classification and scoring at 10 m, the published evidence so far, from a handful of single-region studies, says a few hundred to a few thousand labels on embeddings match or beat a purpose-built imagery pipeline. The independent results count for more than the vendor's.
The AlphaEarth Foundations paper evaluated 15 tasks drawn from 11 open datasets and reports that the embeddings reduced error magnitudes by about 23.9% on average against the next-best approach. It deliberately used weak downstream learners, k-nearest neighbours and linear layers, and tested at one, ten and maximum samples per class. The gains shrink as labels vanish: about 10.4% at ten-shot and 4.18% at one-shot. Those are Google's own numbers on Google's chosen benchmarks, so treat them as a vendor claim.
Independent work is more useful. A USGS-affiliated team published a crop-type comparison in central California for 2020: a random forest trained only on the embedding layer reached 94.7% overall accuracy against the state's reference map, versus 91.9% for a random forest on Landsat and NAIP inputs, and the embedding workflow used three times less cloud compute. Pasture, grain and fallow classes stayed below 65% in both, which tells you the embedding did not fix classes that are ambiguous on the ground.
A September 2026 preprint on old-growth forest detection found that at 10 m, convolutional networks added no benefit over pixel-based XGBoost on the same features. That study did not test a commercial pipeline, but it suggests that at 10 m a heavier vision model adds little once the features are good, so ask what accuracy the bespoke stack in your per-scene contract actually buys.
Prithvi's label efficiency shows up in segmentation. A team that fine-tuned the 300M and 600M models for coastline delineation on small sandy islands reports F1 of 0.94 and IoU of 0.79 with as few as five training images. The Prithvi-EO-2.0 paper reports an 8% improvement over the first Prithvi across GEO-Bench tasks. Both numbers come from the people who built or used the model on one task, so test on yours.
Our earlier analysis of tabular foundation models against tuned XGBoost found the same pattern from the other side: below a modest label count, pretrained representations beat hand-built features, and above it the gap closes.
The Cost Model, in Three Separate Lines
Split the bill into imagery licensing, compute and the model, because the open options zero out two of the three lines while the per-scene contract bundles all of them into one number you cannot audit. Use one defined workload to compare: annual scoring of a 50,000 km² territory at 10 m, with 1,000 to 3,000 labelled points.
Imagery licensing. Sentinel-2 and Sentinel-1 data are free under the Copernicus Sentinel data legal notice, which permits reproduction, adaptation and combination with other data, provided you credit "Copernicus Sentinel data [Year]". Landsat is free from USGS, and HLS from NASA. The AlphaEarth layers themselves are CC-BY 4.0. Commercial imagery is the only line here that costs money.
Compute. At 64 bytes per pixel (the paper quantises the released embeddings to 8 bits), 50,000 km² is 500 million pixels, or about 32 GB per annual layer. Gradient boosting on that runs on CPUs you already own. If you prefer to work inside Earth Engine, note that an insurer or utility is a commercial user: the access guide reserves the free Community Tier for noncommercial projects, and Google's Earth Engine pricing page lists the Basic plan at $500 per month and Professional at $2,000 per month, with included EECU-hour allowances, as checked October 11, 2026. Pulling the COGs from the Source Cooperative mirror avoids that subscription entirely. Prithvi and Clay add GPU time for fine-tuning and inference, which you size against your own throughput; neither publisher prices it for you.
The model. $0 for all three open options. Planet's pricing page did not show a public per-scene or per-hectare rate when checked on October 11, 2026, so treat the incumbent as quote only, and ask for the line items unbundled.
Resale is where buyers get caught. Planet's EULA for Planet Labs BV defines a Derivative Product as one that contains no source image data, says you may redistribute those freely, and requires a licence upgrade to resell imagery or value-added products. If you sell scores to your own customers (a utility's vegetation risk feed to a regulator, an agribusiness selling field reports), confirm which EULA version you signed, because the terms differ by contracting entity. CC-BY 4.0 on AlphaEarth only asks you to carry the attribution line "The AlphaEarth Foundations Satellite Embedding dataset is produced by Google and Google DeepMind."
Where Resolution and Revisit Set a Hard Ceiling
No foundation model rescues a pixel that is too coarse or a revisit that is too slow for your task, so check the input against the question before you compare models. AlphaEarth gives you one 10 m summary per year. That answers "is this parcel cropland, forest or impervious surface this year" and "did it change since last year." It does not answer "did this roof lose shingles in last week's hail" or "is this conductor span encroached by a single tree."
Prithvi inherits HLS: 30 m pixels, with NASA quoting a combined revisit of every 1.6 days once Landsat and Sentinel-2 are merged (1.4 days in 2025 with five satellites), before cloud cover takes its share. That is good for floods, burn scars and crop phenology, poor for anything parcel-sized in a dense suburb.
Clay is the only open option trained on aerial sources such as NAIP and LINZ alongside satellites, so it is the one to test when your labels sit on objects smaller than a 10 m pixel. You still need the high-resolution imagery as input, and that imagery has its own licence.
So the honest split for an insurer is: annual exposure and land-cover features from embeddings, event-driven damage assessment from a commercial vendor.
How to Validate Without Fooling Yourself
Hold out by geography and by season or year, never at random, because neighbouring pixels share information and a random split leaks it into your test set. Ploton and colleagues showed in Nature Communications that large-scale ecological mapping models that looked accurate under random validation performed poorly under spatial validation.
The old-growth preprint shows the same effect on embeddings directly. Under unbuffered spatial validation, the best embedding (Cambridge's TESSERA) beat AlphaEarth by 0.08 PR-AUC. With a 10 km buffer between training and test locations, the margin fell to 0.03, with an interval "consistent with no difference." The ranking you would have published depended on the buffer.
For annual embeddings, add a year holdout. The Henan winter wheat study trained on labels from 2020 and mapped 2018 to 2024, reaching 0.85 overall accuracy with cosine similarity and linear regression and 0.86 to 0.93 with lightweight classifiers. That is the test that matches production: labels from one year, predictions on the next.
Set the buffer to at least the distance over which your target is correlated (a field, a feeder circuit, a watershed) and report the number from that split to whoever signs off the renewal.
The Embedding Space Will Change Under You
Pin one dataset or weights version, record it next to every prediction, and never train a downstream model on features from two vintages. Each of the three has already changed once.
AlphaEarth. The catalog says the current layers carry DATASET_VERSION 1.1 and were generated with v2.1 of the model, including "a regenerated 2017 layer that incorporates additional Sentinel-1 acquisitions." Google states that "The embedding space is also consistent across years" and that it is "committed to ongoing production of annual Satellite Embedding layers," with at least a year's notice of changes to delivery. Neither sentence is a contract. If your pricing model depends on the 2027 layer landing in the same space as 2024, get that promise in writing from Google or plan to re-fit every year on a single vintage.
Prithvi. Prithvi-EO-1.0 was a 100M model pretrained on contiguous-US HLS data. Version 2.0 is a different set of global models at 300M and 600M parameters. A fine-tuned head from 1.0 does not transfer.
Clay. v1.5 is a separately trained 632M model, and the project does not publish a compatibility statement against v1.0 vectors. Clay has also changed stewards: the Source Cooperative listing says v0 through v1.5 were a fiscally sponsored project of Radiant Earth Foundation, and it is now a Renaissance Philanthropy program. That is why Clay loses this comparison for a production buyer. The model is capable, but you are betting on a nonprofit's roadmap with no stability commitment. Mirror the weights you use, as we argued when Hugging Face hired bankers.
We covered the same failure in text retrieval, where re-embedding 10M chunks after a model swap became a line item nobody had budgeted. Satellite embeddings carry the same risk, plus a regulator who may ask you to reproduce last year's score.
Who Should Not Pick Each Option
Each option has a buyer it will fail, and that buyer should find out before the pilot.
Skip AlphaEarth if your answer has to be fresher than the last completed year, finer than 10 m, or reproducible from a model you can run yourself. You cannot fine-tune it; you only get its outputs.
Skip Prithvi if you have no one to own a GPU training loop and a serving endpoint, or if your objects are smaller than a 30 m pixel.
Skip Clay if your compliance team needs a named commercial steward and a version policy. Pick it only when your labels depend on aerial imagery the other two never saw.
Skip renewing the per-scene contract for annual 10 m scoring. Keep it, narrowed, for daily monitoring and sub-metre damage work. Insurers should also check how a third-party model fits the model inventory regulators are starting to expect, which we covered in the NAIC's draft AI exam.
What Predicts Regret, and What You Do Next
The decision turns on three questions: whether the task survives a 10 m annual pixel, whether you can label a few hundred spatially separated points, and whether you can live with a version pin. If all three are yes, the per-scene contract is overpriced for that job. If the first is no, the open models are the wrong comparison.
This Week: Pull the current contract and split its invoice into imagery, compute and analytics. Download one year of AlphaEarth tiles for your territory from Source Cooperative and record the dataset version.
This Month: Fit XGBoost on your existing labels with a geographic buffer holdout and a year holdout. Compare against the vendor's scores on the same held-out parcels. If you need segmentation, run the same test with Prithvi-EO-2.0-300M.
Before the Renewal Date: Ask Google, in writing, what it commits to on embedding-space stability for future annual layers, and ask your imagery vendor for the Derivative Product clause in the EULA you signed. Narrow the renewal to the daily or sub-metre work the embeddings cannot do, and quote the buffered accuracy number in the business case.
Continue Reading
- Foundation Models Beat Tuned XGBoost Below 50,000 Rows
- Mitra vs TabPFN vs XGBoost: Free Mitra Beats Trees on Small Tables
- OpenAI vs Cohere vs Qwen3: Re-Embedding 10M Chunks Costs $520
- Hugging Face Hired Bankers. Go Mirror Your Weights.
- TimesFM vs Chronos vs Moirai: The Leader's Weights Ban Production
- Scale AI Alternatives: Run Two Vendors and Keep the Gold Set
- NAIC's Draft AI Exam Lets Insurers Set Their Own Materiality Bar
