Reflection's Beam Beats Kimi K3 on One of 14 Shared Benchmarks

Reflection's Beam is a 501B Apache 2.0 open model from a US lab, and its own table has it losing to Kimi K3 on 13 of 14 shared benchmarks. Weights and pricing aren't out yet.

By Rajesh Beri·October 6, 2026·10 min read
Share:
A server rack in a locked government data centre with one empty, unlit bay and a printed shipping label taped to the empty slot, beside a lit rack full of running servers.

Illustration generated using AI

If your policy bars Chinese-origin weights, Reflection AI's Beam is a new US-built Apache 2.0 option, and by Reflection's own numbers it sits a full tier below the Chinese model you can't use. In the benchmark table Reflection published on October 5, Beam loses to Moonshot's Kimi K3 on 13 of the 14 rows where both have a score. The weights aren't out yet and there is no price. The useful move this month is to build the evaluation that will decide whether Beam replaces anything in your stack, so you can run it the day the weights land.

Beam is a 501-billion-parameter sparse mixture-of-experts model with 23 billion parameters active per token, a 256K pretraining context extended to 1M tokens in midtraining, and a promised Apache 2.0 license, according to Reflection's launch post. Reflection says the model is "undergoing final red-teaming and evaluations," with "weights, technical report, model card, and developer artifacts later this month." Today you get a sign-up form for early access at platform.reflection.ai.

What Did Reflection Actually Ship on October 5?

Reflection shipped an announcement, a benchmark table and a waitlist; the downloadable model is still a promise. Implicator's write-up lists what is missing as of launch day: no weights, no technical report, no pricing, no independent verification and no compressed builds for smaller hardware.

What Reflection did disclose is unusually detailed on training. Beam was pretrained on 23.8 trillion tokens in under four weeks on 6,144 NVIDIA GB300 GPUs, then put through reinforcement learning on 10,500 GB300s for another four weeks, generating more than 100 million rollouts across roughly a million environments, per the launch post. Latent Space's summary adds that part of the pretraining corpus came from OCR over hundreds of millions of PDFs, and that early analyst readings place Beam around the GLM-5.2 tier and below DeepSeek V4 Flash on some benchmarks.

The company behind it is well funded and explicit about how it plans to make money. Reflection raised $2 billion at an $8 billion valuation in October 2025, and told TechCrunch then that weights would be public while datasets and training pipelines stay proprietary. Revenue is meant to come from large enterprises building on the models and from governments commissioning sovereign systems. That matters for your budget: the weights will be free, and the support contract, the hosted endpoint and any sovereign deployment are where Reflection will put a price.

Where Does Beam Trail Kimi K3, and by How Much?

Beam trails Kimi K3 on every coding, search and long-context row the two share, and beats it only on tau3 banking, by 0.9 points. These are Reflection's numbers, from Reflection's table; nobody outside the company has reproduced them yet.

Benchmark (Reflection's table) Beam Kimi K3 Gap
SWE Atlas Codebase QnA 34.6 68.0 -33.4
DeepSWE v1.1 44.4 68.0 -23.6
DeepSearchQA (context management) 80.1 95.0 -14.9
BrowseComp (context management) 77.4 91.2 -13.8
SWE-Bench Pro v2-Hard 77.2 88.2 -11.0
HLE, no tools 36.2 46.9 -10.7
AutomationBench public 37.0 46.7 -9.7
AA-LCR 79.3 88.7 -9.4
SciCode 49.7 58.7 -9.0
Terminal-Bench v2.1 80.1 88.3 -8.2
CritPt 16.3 23.4 -7.1
MCP Atlas 78.7 82.3 -3.6
GPQA Diamond 90.5 93.5 -3.0
tau3 banking 38.0 37.1 +0.9

Source: Reflection, "Introducing Beam," October 5, 2026. Gaps are simple subtraction.

The biggest gaps sit on the work enterprises are actually buying open models for. SWE Atlas Codebase QnA asks a model to answer questions about a real codebase, and Beam scores about half of Kimi K3's result. DeepSWE measures agentic software engineering, where Beam lands at 44.4 against 68.0. The rows where Beam comes close, GPQA Diamond and MCP Atlas, are science questions and tool calling, much shorter tasks than a multi-step coding run.

The steel-man for Beam is size. Kimi K3 is a 2.8-trillion-parameter model with 104 billion active per token, per Digital Applied's read of the release, so Beam is roughly a fifth of its total size with under a quarter of the active compute. Reflection's own pitch is efficiency: Beam "achieves scores comparable to GLM-5.2 while using 3–4× less inference compute." Its table backs the GLM-5.2 comparison reasonably well (44.4 vs 44.0 on DeepSWE, 80.1 vs 81.0 on Terminal-Bench, 38.0 vs 37.1 on tau3 banking). Against Inkling and NVIDIA's Nemotron 3 Ultra, the other non-Chinese-lab entries in the table, Beam wins most shared rows, 80.1 to 56.4 against Nemotron 3 Ultra on Terminal-Bench, though Nemotron 3 Ultra edges it on IFBench (81.7 to 79.7) and ties it on AA-LCR.

If your policy leaves you only non-Chinese open models, Beam is the strongest entry in Reflection's table. Against Kimi K3, it loses 13 of 14.


Why Would a Buyer Accept a Weaker Model?

A buyer accepts it because provenance rules decide the shortlist before benchmarks get a vote. A bipartisan group of US lawmakers introduced the No Adversarial AI Act in June 2025 to bar executive agencies from AI systems developed in China, Russia, Iran and North Korea, with the Federal Acquisition Security Council maintaining the list. If your company has copied a rule like that into its own model policy, Beam is aimed squarely at you.

The security case behind those rules has evidence attached. A CNAS fellow's August 2026 analysis in Just Security notes that the US government's CAISI has published five assessments of Chinese open-weight models since January 2025, and that a DeepSeek model complied with every hacking and scam request CAISI tested while American models refused nearly all of them. The same piece argues against blanket bans, since downloaded weights cannot be recalled, and proposes testing standards instead.

The license is the second reason. Kimi K3 ships under a bespoke license, tagged license:other on Hugging Face: a model-as-a-service operator whose group revenue tops $20 million over any 12 months must sign a separate agreement with Moonshot, and products past 100 million monthly users or $20 million in monthly revenue must display "Kimi K3," per Digital Applied. Internal use is exempt. Apache 2.0 has no revenue trigger, and its Section 3 grants a perpetual, royalty-free patent license from every contributor, which ends only if you sue claiming the work infringes your patents. If you resell an AI product, Apache removes a negotiation with Moonshot that Kimi K3 would require past that revenue line.

If you sell software built on an open model, a vendor-supplied model provenance trail is now a procurement question. We covered how that played out in legal AI in CoCounsel's New Model Runs on Qwen. Go Read the Card.

What Will Beam Cost to Run?

Nobody can price Beam yet, because Reflection has published no API price and no checkpoint precision. You can still size the hardware.

A mixture-of-experts model loads all of its experts into memory even though only some fire on each token; we worked through that trap in Tencent Says 49B Active. Your Node Loads All 770B. For Beam, 501 billion parameters at one byte each (FP8) is about 500GB of weights before the KV cache; at 4-bit it is roughly 250GB. That is arithmetic on Reflection's parameter count, and the real figure depends on the precision Reflection ships. For comparison, Kimi K3 needs about 1.4TB resident at its native MXFP4 before any KV cache, with a 1.56TB download, per Digital Applied. If Beam ships at FP8, it plausibly fits on a single node of eight 80GB GPUs, while Kimi K3 needs nearly three times the memory. That is the strongest operational argument Reflection has.

Treat the "3–4× less inference compute" claim with care. Implicator reports that Reflection calls the comparison approximate rather than measured inference cost, excluding prompt processing, attention and serving overhead. On a long-context agent workload, prompt processing is often most of the bill.

What Could Change Before the Weights Ship?

The release date, the license scope and the benchmark numbers can all still move, and Reflection has said so. The company told The Next Web that US and UK government AI institutes are assessing the model with it. Co-founder Ioannis Antonoglou also said: "It is possible that you get to a level of capability that you want to just be more careful with how you deploy it." That is a general statement about frontier capability, but it is a reason not to build a Q4 plan on a download date.

The technical report is the other gate. Reflection's competitor columns contain many "NR" cells, and the post says only that competitor scores come from Artificial Analysis and DataCurve. Until the report explains the harness, scaffolding and sampling settings, treat every row as a vendor claim. Hugging Face co-founder Clem Delangue's challenge to Reflection from October 2025, quoted by Implicator, still applies: the test is "high velocity of sharing of open AI models and datasets." We have seen a lab quietly swap a model under a stable name before; see DeepSeek Swapped the Model. Your Eval Didn't Notice.


What Should You Do Before the Weights Drop?

Build the evaluation now, decide what Beam would replace, and wait for the technical report before signing anything.

This Week:

  1. Get your model policy in writing from security and legal: does it bar Chinese-origin weights outright, only Chinese-hosted APIs, or only certain licenses? Beam's case rests entirely on that answer. If the policy only bars hosted APIs, self-hosted Kimi K3 may already be allowed, and Beam is a weaker substitute.
  2. Pick the one workload Beam would take over (a closed-API coding assistant, an air-gapped RAG service, a restricted Chinese model in a sandbox) and write down the model it would displace.
  3. Join the early access waitlist from a work account so your team gets the API before the weights are public.

This Month:

  1. Assemble 100 to 200 tasks from your own repositories and tickets, scored the way your engineers would score them. Public rows like SWE Atlas and DeepSWE tell you where to look; your tasks tell you whether a 24-to-33-point gap matters for your work.
  2. Run that set today against the model you would replace and against Nemotron 3 Ultra or whichever non-Chinese open model you already host, so Beam has a baseline to beat on day one.
  3. Reserve capacity for a single eight-GPU node for a two-week bake-off, sized to the FP8 arithmetic above, and set up a private mirror so the weights you test are the weights you deploy. The playbook is in Hugging Face Hired Bankers. Go Mirror Your Weights.

Before You Commit:

  1. Read the technical report for the evaluation harness, and rerun at least two of Reflection's rows (Terminal-Bench and SWE-Bench Pro are public) on your own hardware.
  2. Confirm the license file on the published checkpoint actually says Apache 2.0, with no added acceptable-use addendum.
  3. Ask Reflection in writing what a support contract costs and what it covers. Free weights do not come with a security patch process.

The Bottom Line

Beam gives policy-restricted enterprises a credible, permissively licensed open model at roughly GLM-5.2 quality, running on hardware they can plausibly afford. It does not give them Kimi K3's capability, and Reflection's own table says so. The air-gapped stack we recommended in Best Air-Gapped LLM Stack: Apache Weights on vLLM Beat NVIDIA's Fee gets a stronger Apache option, and IBM's choice of Nemotron and Laguna for air-gapped Bob now has a new comparison point.

Have your 200-task eval ready before Reflection posts the checkpoint, and run it the same day.

Continue Reading

Share:

Frequently Asked Questions

What is Reflection AI's Beam model?

Beam is a 501-billion-parameter sparse mixture-of-experts language model with 23 billion parameters active per token and up to 1M tokens of context, announced by Reflection AI on October 5, 2026. Reflection says it will release the weights under Apache 2.0 later in October.

How does Reflection Beam compare with Kimi K3?

In Reflection's own table, Beam trails Kimi K3 on 13 of the 14 benchmarks where both have a score, by 33.4 points on SWE Atlas Codebase QnA and 23.6 on DeepSWE. It beats Kimi K3 only on tau3 banking, 38.0 to 37.1. No independent lab has reproduced these numbers yet.

Can I download Reflection Beam weights yet?

Not as of October 6, 2026. Reflection says the model is undergoing final red-teaming and that weights, a technical report and a model card will follow later in October. Early access is a sign-up at platform.reflection.ai, and no price has been published.

How much GPU memory will Beam need to self-host?

Reflection hasn't said what precision it will ship. By simple arithmetic, 501 billion parameters is about 500GB of weights at FP8 or about 250GB at 4-bit, before KV cache. Kimi K3 needs about 1.4TB resident at its native MXFP4.

Why choose Beam over a stronger Chinese open model?

Mainly policy and licensing. Some organizations bar Chinese-origin models, and Kimi K3's custom license requires a separate Moonshot agreement for model-as-a-service operators above $20 million in revenue. Apache 2.0 has no revenue trigger and includes a royalty-free patent grant.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →