The DOE Wants Your Lab Data. You Write the Contract.

DOE's Genesis Open Models portal wants proprietary scientific data and the first window closes August 14. Applying transfers nothing — but there is no published contributor agreement, so the handling terms are yours to draft.

By Rajesh Beri·August 9, 2026·13 min read
Share:
A sealed cardboard archive box of laboratory notebooks and printed spectra charts sitting alone on a concrete loading dock outside a low federal laboratory building at dusk, roller door half open, blank shipping label on

Illustration generated using AI

The Department of Energy is asking private companies to hand over proprietary scientific data for a federal open-weight model, and the first window closes on August 14 — five days from now. The application itself is free and moves nothing. The decision that matters comes after it, and almost nobody has drafted for it: there is no published contract, so whatever protection your data gets is protection you write yourself.

DOE launched the Genesis Open Models Initiative on August 7, with a contribution portal hosted at Argonne National Laboratory. It wants domain-specific scientific data, open-weight base models, fine-tuned variants, and evaluations, for a family of models aimed at materials discovery, energy systems, earth modeling, fusion, biology and high-energy physics. If you run R&D at a pharma, chemicals, materials, semiconductor or industrial company, you hold exactly the corpora being asked for.

Here is the part the coverage keeps burying. The public application form collects descriptions and metadata only — no scientific material transfers at the application stage — and each contributor specifies how its own material may be handled. That single sentence is the entire commercial terms of this program. You are not accepting a licence. You are proposing one.


What DOE Is Actually Asking Companies to Contribute

DOE is soliciting three distinct things, and they carry very different risk. Per the department's own announcement, the tracks are: open-weight models with transparent provenance and documented training procedures; pretraining contributions, meaning domain-specific datasets, benchmarks and specialised corpora; and fine-tuning contributions, meaning domain-adapted model versions.

The pretraining track is the one with teeth. Arcee's description of the program lists foundation-stage contributions as scientific text, code, documentation and structured collections suitable for pretraining, midtraining or context extension. Post-training contributions are the longer list: expert demonstrations, annotated task examples, research software, datasets, complete workflow environments, reinforcement-learning tasks, held-out evaluations, scoring rubrics, tests and verifiers.

Read that list against your own asset register. Assay results, characterisation data, failed-synthesis records, internal simulation corpora, instrument logs, the annotated notebooks your senior scientists built over fifteen years — that is what "high-quality, domain-specific scientific data" means in practice. Your negative results are, for a scientific model, among the most valuable things you own, and they are the things you have never bothered to classify because nobody ever asked for them.

The Application Is Free. The Gate After It Is Not.

Applying costs you a form and transfers nothing, which is why the deadline is a bad thing to agonise over. What you should agonise over is the second review gate.

Applications run through five review gates: scientific fit, rights and handling, expert and evaluation readiness, technical integration, and final program selection. Gate two is where your legal position gets tested — and where, if you have not already decided what you will and will not permit, you will either stall out or agree to something under time pressure.

Applicants are asked to identify their track, describe the scientific work and why it matters, explain what material is ready, name the experts available to support it, and state the proposed terms of use. That last field is not boilerplate. It is the negotiation, compressed into a text box, submitted before anyone from your legal team has seen a counterparty agreement — because there isn't one to see.

What contributors get in return is specific and modest: early evaluation access during system development, and credit in the technical report and release materials. No cash. No royalty. No equity. If your board asks what the consideration is, the honest answer is early access and a citation.

The Executive Order Promised a Standard Agreement. The Portal Doesn't Have One.

This is the gap that should shape how you respond, and it is visible in the primary text.

Executive Order 14363, signed November 24, 2025, directs DOE to build the American Science and Security Platform with "secure access to appropriate datasets, including proprietary, federally curated, and open scientific datasets." It also directs the Secretary to "develop standardized partnership frameworks, including cooperative research and development" agreements, with "clear policies for ownership, licensing, trade-secret protections, and commercialization of intellectual property developed under the Mission."

Standardised frameworks. Trade-secret protections. Eight months later, the contribution portal asks you to state your own terms.

That is not necessarily bad faith — building a standard data-use agreement across 17 national laboratories is slow, and the program is moving faster than the paperwork. But it changes what you are doing. Outside counsel reading the same EO flagged the missing pieces early, noting DOE must develop "standardized cooperative research, data-use, and model-sharing agreements" and "uniform data-access and cybersecurity standards for non-federal collaborators." Another firm's guidance to in-house counsel is blunter about what those agreements need to nail down: rights in contributed datasets, "including confidentiality, intellectual property ownership, derivative model rights, and limits on redistribution."

Derivative model rights and limits on redistribution are the whole ballgame, and in an open-weight program they are structurally in tension. The output is meant to be downloadable by anyone, including the competitor down the road who contributed nothing.

There is a well-worn federal instrument that already solves part of this, and you should ask for it by name. Under a Cooperative Research and Development Agreement, information produced in performance of the CRADA can be marked Protected CRADA Information and withheld from public disclosure for up to five years — a defined statutory shelter rather than a promise. It is not a perfect fit for a data-contribution portal. It is a far better starting draft than a text box.


The Model You Would Be Contributing To Does Not Exist Yet

Steel-man the program first, because the case for it is real. Open weights that an institution can hold and operate directly are the only version of frontier capability that a national lab, a hospital system or a regulated manufacturer can actually deploy without exporting its data to somebody else's API — the same argument that makes open-weight models a genuine hedge against API pricing, and the same logic behind the federal move toward open-source AI in defence. DOE's CEO partner puts the national case plainly: "A country cannot lead in AI if everything it leads in is closed," said Arcee AI's Mark McQuade.

Now the ledger. As of the launch announcement, Genesis-Science-1 has no published parameter count, no benchmarks and no released weights. Neither the DOE page nor Arcee's own page names a licence. You are being asked to commit proprietary data to an artifact whose size, quality and distribution terms are all unspecified.

The best available proxy is what the development partner shipped last. Arcee's Trinity Large is a 400-billion-parameter sparse mixture-of-experts model with 13 billion active parameters per token, trained on 17 trillion tokens across 2,048 NVIDIA B300 GPUs in 33 days, released January 27, 2026. That is a serious engineering result from a small lab. Note the second half of the same page: the dataset was curated by a third party and included over 8 trillion tokens of synthetic data, and the weights were published — the training data was not.

Apply that pattern to your contribution and you get the shape of the deal. Your data does not become public. The model trained on it does. Anyone can download it, including your competitors, and nothing in the credit line stops them.

What Your Data Actually Becomes

Once a corpus is folded into a set of open weights, it stops being a dataset and becomes a capability that ships. Three consequences follow, none of them hypothetical.

Attribution disappears. A blended pretraining corpus does not carry per-contributor provenance downstream. The technical report credits you; the weights do not, and neither does the fine-tune somebody derives from them six months later.

Redistribution is the point, not a risk. In a proprietary licensing deal you can cap the field of use — that is exactly the structure enterprises have used when selling data into the AI supply chain. An open-weight release has no field of use. Whatever your terms say about your data, the derived model is meant to be free to run.

Extraction is a live research area, not a settled one. A 2026 evaluation of memorisation in open models found that, for one of the two models it tested, adversarial prefix attacks produced a 36-fold increase in near-verbatim recall of training data compared with ordinary prompting — the authors' own conclusion being that models "can reveal training data when prompted adversarially, but they rarely do so in more common prompting conditions." That does not mean your assay records will fall out of Genesis-Science-1. It means "the data itself is never released" and "the data can never be recovered" are different claims, and only the first one is being made.

None of this argues for staying out. It argues for knowing which slice of your corpus you would put in. The answer for most companies is the same shape: published-but-unstructured material, retired program data, characterisation records past their competitive half-life, and negative results — genuinely useful for training, genuinely not the crown jewels. If your data classification cannot tell those apart today, that is the real finding, and it is the same data-readiness gap that stalls internal AI programs.

Pharma Already Solved This a Different Way

The last time an industry with genuinely irreplaceable proprietary data tried to build a shared model, it did not move the data at all.

MELLODDY brought ten pharmaceutical companies — Bayer, GSK, Novartis, AstraZeneca, Amgen and others — into a single drug-discovery model using federated learning. Each partner trained locally and contributed only securely aggregated gradient updates, so that the "private underlying data and resulting head models never leave the respective owner-controlled architectures, in any form" — and each of the ten realised aggregate improvements on its own classification or regression models. The gains were real but not uniform: they concentrated in pharmacokinetics and safety-panel tasks, and on regression all but one partner beat their single-partner baseline. Direct competitors got better models without a single compound structure leaving anyone's infrastructure.

That is the comparison worth putting in front of your CTO. It proves the collaborative benefit is real, and it proves centralised contribution is not the only architecture that delivers it. Genesis is asking for the centralised version. It is entirely legitimate to apply and propose the federated one — a compute-and-expertise contribution, an on-premises fine-tune delivered as weights rather than as data, or evaluation sets and verifiers instead of a corpus. The post-training track explicitly wants rubrics, held-out evaluations and verifiers, and those leak far less than a pretraining corpus while making you materially harder to remove from the program.

Why the Window Is Only Five Days Wide

The compression is not about your readiness. It is about a clock that started in November.

EO 14363 gives DOE 270 days to demonstrate initial operating capability of the Platform for at least one identified challenge. Counting from the November 24, 2025 signing date, that lands on August 21, 2026 — one week after the pretraining application deadline. The 120-day milestone required a plan for incorporating datasets from "approved private-sector partners." The portal is that plan, arriving late and running hot.

Two things follow. First, the deadline pressure is the program's, not yours: DOE has said additional windows will be rolling, expected every three months. Missing August 14 costs you a quarter, not the program. Second, the published dates do not agree with each other. DOE's August 7 page says pretraining applications close August 14 and fine-tuning August 25. Arcee's July 23 release says foundation-stage contributors apply by August 6 and deliver by August 20, with post-training applications by August 25 and delivery by September 14. Third-party coverage lists apply by August 14 and deliver by August 28. Treat the DOE page as authoritative — it is the agency's own and the most recent — and get your specific window confirmed in writing before you plan any delivery around it.

For context on who is already inside the tent: DOE announced collaboration agreements with 24 organisations in December 2025, including AWS, Google, Microsoft, NVIDIA, AMD, Intel, IBM, Oracle, Anthropic, OpenAI, xAI, Palantir, Cerebras, Groq and CoreWeave. The compute and model vendors signed up eight months ago. The data holders are being asked now.

What to Do Before August 14

This Week (before August 14): Have your R&D and legal leads answer one question on one page — which specific dataset would we offer, and under what handling terms? If you can answer it, file the application; it transfers nothing and buys you early evaluation access and a seat at gate two. If you cannot, do not file. Skip to the next quarterly window with an answer ready. Also assign someone to email the Argonne-hosted portal and get your track's actual application and delivery dates confirmed in writing, given the three conflicting published sets.

Before You Submit Anything Beyond Metadata: Draft your proposed terms of use as a document, not a text box. Minimum contents: permitted use limited to training the named model; no redistribution of the underlying data; a defined confidentiality period; explicit treatment of derivative model rights; and named export-control and foreign-national-access handling. Ask directly whether DOE will paper the contribution as a CRADA with Protected CRADA Information marking and its statutory withholding period. If the answer is no, you have learned the most important fact about this program before it cost you anything.

This Month: Classify the R&D corpus you would actually contribute, at file-set granularity, into three tiers — contributable, contributable-under-terms, never. Most organisations discover the exercise is worth doing regardless of Genesis, because the same tiers govern every vendor evaluation and every internal fine-tune you will run this year.

This Quarter: Put the federated alternative on the table as your counter-proposal. Offer evaluations, verifiers, expert time, an on-premises fine-tune, or compute — contributions that earn you standing in the program without moving a corpus. If DOE will only take raw data, that tells you something about the program's maturity, and you will have learned it from a proposal rather than from a breach.

The Bottom Line

Federal AI programs have been reshaping enterprise vendor risk all year — from pre-deployment testing of frontier models to the government taking direct positions in AI vendors. Genesis inverts the flow. This time the government is not buying, regulating or investing. It is asking you to contribute the one asset that is genuinely hard to reproduce.

That may well be worth doing. A credible American open-weight scientific model, held and operated by the institutions that need it, is a public good your own researchers would use — and licence clarity is exactly what made Apache-2.0 open weights usable in the enterprise in the first place. But the deadline is not the decision. The decision is what you are willing to put in and on what terms, and right now those terms are yours to write, which means they are also yours to get wrong.

Do not send data because a date is close. Send terms because you drafted them.

Continue Reading

The Pentagon Went Open-Source AI. Your Lock-In Excuse Just Died. Your AI Vendor's New Boss: Washington's $42.6B OpenAI Stake D&B Signed With All 3 AI Giants in 4 Weeks. Data Is the New Moat. Only 5% Are AI-Ready: The 25-Point Scorecard Federal AI Vendor Risk: The $32B Test Every CIO Must Run DeepSeek Will Raise Prices. Your Ceiling Is Already 4x. Google Gemma 4: Why Apache 2.0 Changes Enterprise AI

Share:

Frequently Asked Questions

What is the DOE Genesis Open Models Initiative?

A Department of Energy program launched August 7, 2026 to build open-weight foundation models for science, starting with Genesis-Science-1, developed with Arcee AI and hosted through a contribution portal at Argonne National Laboratory. It solicits open-weight models, domain-specific pretraining data, fine-tuned variants, evaluations and expertise across materials, energy, earth modeling, fusion, biology and high-energy physics.

When is the Genesis Open Models contribution deadline?

DOE's August 7 announcement says pretraining contribution applications close August 14, 2026 and fine-tuning applications close August 25, 2026, with additional windows rolling roughly every three months. Arcee's earlier July 23 release lists different dates, so confirm your specific track's application and delivery dates in writing before planning around them.

Does applying to Genesis Open Models require handing over data?

No. The public application form collects descriptions and metadata only — no scientific material transfers at the application stage, and each contributor states its own proposed terms of use. Applications then pass through five review gates, the second of which covers rights and handling. That gate, not the application, is where your legal exposure is decided.

What do contributors get in return for their scientific data?

Early evaluation access during system development, and credit in the technical report and release materials. There is no cash payment, royalty or equity described in the program materials, and as of launch Genesis-Science-1 has no published parameter count, benchmarks, released weights or licence.

How can a company protect proprietary data contributed to a DOE program?

Ask whether DOE will paper the contribution as a Cooperative Research and Development Agreement. Information produced under a CRADA can be marked Protected CRADA Information and withheld from public disclosure for up to five years — a defined statutory shelter rather than a self-drafted promise. Your proposed terms should also cap permitted use to the named model, bar redistribution of the underlying data, and address derivative model rights explicitly.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →