Self-hosted IBM Bob runs one local model, NVIDIA Nemotron or Poolside Laguna, in place of the Claude, Mistral and Granite mix the SaaS version routes between. If you block cloud coding agents because source code cannot leave your network, that is the same agent with a different brain. IBM has published no evals of Bob on either local model, no GPU sizing and no price. Treat it as a new product that shares a user interface with the old one, and run your own bake-off on your own repositories before you sign.
IBM announced self-hosted deployment for Bob on October 1, covering on-premises, private cloud, sovereign cloud and air-gapped environments, plus hybrid configurations that connect to external model services. IBM shares rose more than 3% in late trading after the news. The pitch lands on a real constraint: IBM's own Institute for Business Value found 68% of executives find data residency and sovereignty requirements challenging. Banks, agencies and insurers that said no to Claude Code and GitHub Copilot over code egress now have an IBM-branded yes on the table, and the details of that yes sit in IBM's installation docs rather than its press release.
What Self-Hosted Bob Actually Runs
Self-hosted Bob runs a single core model, and IBM's announcement names two for customer-managed infrastructure: nvidia/nemotron-3-ultra-550b and poolside/laguna-s-2.1. Both IDs appear in IBM's model gateway configuration docs, which also list a Mistral Medium ID among the configurable core models (the announcement does not mention it, so ask IBM whether it is supported) and carry the sentence that matters most for a buyer: "Configure only one core inference model in the model gateway at a time. Configuring multiple core models simultaneously is not supported."
Compare that with how IBM sold Bob in April. The SaaS launch said Bob draws on "Anthropic Claude, Mistral open source models, and IBM Granite, alongside specialized fine-tuned models" and that it "dynamically routes tasks to a suitable model based on accuracy, performance, and cost." Simple completions went to light models, hard tasks to capable ones. The numbers IBM quotes come from that routed setup: 80,000+ IBM employees on Bob and an average self-reported 45% productivity gain, and the 9-months-to-3-days modernization claim we covered in July.
None of those results were measured on a single Nemotron or Laguna deployment. IBM's self-hosted announcement promises "a consistent core Bob experience across supported deployment configurations." That describes the IDE, BobShell, parallel tool calling, skills and modes. It says nothing about output quality.
Two smaller gaps also show up in the docs. Neither Laguna S 2.1 nor Nemotron 3 supports image inputs, so any workflow where a developer pastes a screenshot of a UI bug or an architecture diagram stops working. And the self-hosted overview states that "security event logging and monitoring for Bob self-hosted are managed at the OpenShift platform level and are not provided by Bob." Your SOC has to build it.
How Nemotron and Laguna Compare With the Models You Know
On their makers' own benchmarks, the smaller model is the stronger agentic coder, and both trail the closed frontier. Here is what NVIDIA and Poolside publish, with the caveat that every number below is vendor-reported and none was run inside Bob's harness.
| Nemotron 3 Ultra | Laguna S 2.1 | |
|---|---|---|
| Total / active parameters | 550B / 55B | 118B / 8B |
| Terminal-Bench 2.1 | 56.4 | 70.2% (thinking on), 60.4% (off) |
| SWE-Bench | 70.7 (Verified) | 59.4% (Pro), 78.5% (Multilingual) |
| Minimum hardware | 8x B200 or 8x H200, or 16x H100 | One DGX Spark |
| License | OpenMDW 1.1 | OpenMDW 1.1 |
Sources: NVIDIA's model card and Poolside's launch post.
Three things follow. Laguna scores about 14 points higher than Nemotron on Terminal-Bench 2.1 while running on a fraction of the hardware, so the bigger model is the wrong default if agentic coding is the job. Laguna's lead depends on max thinking, which means longer and more variable response times than the table suggests. And Laguna still sits roughly 10 to 15 points behind the closed systems from OpenAI and Anthropic on Terminal-Bench, according to Poolside's own results as reported by TNW at launch.
The hardware row deserves the CFO's attention. Nemotron's BF16 card lists 8x GB200, B200, GB300 or B300, 16x H100, or 8x H200 as the minimum, because a mixture-of-experts model has to keep all 550B parameters in memory even though only 55B fire per token. We walked through that trap in Tencent's 49B-active, 770B-loaded model. Laguna, by contrast, runs on a single NVIDIA DGX Spark, though a DGX Spark serving one developer is a different sizing problem from a cluster serving 2,000.
There is also a supplier question on the Laguna side. In August we covered NVIDIA hiring 109 Poolside engineers as a continuity risk. Open weights under OpenMDW mean you keep the model you have. Fixes and the next version depend on Poolside's roadmap.
Where the Hybrid Mode Leaks
Hybrid is where IBM's sovereignty pitch is weakest, because the moment Bob calls an external model your code goes to that model's provider. IBM's self-hosted announcement lists the hybrid options as Claude Sonnet 5.0, Claude Opus 4.8, Gemini 3.7 Flash and OpenAI GPT 5.6 Sol, and the gateway docs wire them through OpenAI-compatible endpoints, AWS Bedrock or Google Vertex AI.
Hybrid has a legitimate use. A bank could keep its core payments code on a local Laguna instance and let a lower-risk internal tools team use Claude through Bedrock. IBM describes it as a way to "connect Bob to an approved external model service for a less restricted workload."
The docs do not explain how that squares with the one-core-model limit. If one gateway serves one core model, splitting restricted and unrestricted work probably means separate Bob instances, separate namespaces and a policy that decides which repositories each one can see. That is an inference from the docs, not something IBM has stated, and it is the first question to put to your IBM account team. Ask to see the configuration that guarantees a restricted repository can never be sent to the external endpoint, including by a developer who switches modes.
If that answer is a written policy with no technical enforcement behind it, assume code will leak through it eventually. Our review of coding assistants that survive security review found the same pattern across vendors: security teams approved the tools whose data path was enforced in configuration and could be checked.
What Nobody Has Priced Yet
Self-hosted Bob has no public price, and the full cost has three parts IBM has not added up for you. SaaS Bob lists Pro at $20 a month, Pro+ at $60 and Ultra at $200, each with a Bobcoin allowance; Enterprise is custom. Neither SiliconANGLE nor Startup Fortune reports a price for the self-hosted tier.
The three parts:
- The Bob license. IBM says self-hosted supports bring-your-own-license for eligible existing model investments, and the Java modernization, IBM i and IBM Z packages are premium add-ons "subject to applicable licensing." Expect a negotiated number.
- The GPUs. For Nemotron, that is an 8-GPU Blackwell or H200 node at minimum before you add concurrency. For Laguna it is far less per instance, but you still need enough capacity for every developer who hits the agent at 10 a.m.
- The platform. Self-hosted Bob requires a Red Hat OpenShift cluster, and you own upgrades, scaling, availability, networking, identity and security logging. If you already run OpenShift, this is marginal. If you do not, you are buying and staffing a container platform along with the coding agent.
The steel-man for IBM is real. A regulated buyer with no permitted cloud option has had few agentic choices, and IBM is shipping one with a support contract, an operator and Helm charts. Our air-gapped LLM stack comparison found that assembling this yourself is doable and cheap on licenses, but you then own the agent harness too. Bob saves you that. The question is whether it saves enough to justify a price IBM has not named.
What to Do Before You Sign
The decision rests on one measurement IBM has not published: how well Bob performs on your code with the local model you would actually run. Get that number yourself.
This Week:
- Email your IBM account team three questions in writing: the self-hosted license price per developer, the GPU sizing IBM recommends per 100 concurrent developers for each supported model, and whether IBM has eval results for Bob on Nemotron and Laguna versus SaaS Bob.
- Ask how a hybrid deployment technically prevents a restricted repository from reaching an external model, given the one-core-model limit. Ask for the actual gateway and namespace configuration in writing.
- Pick 20 closed tickets from your own backlog (a mix of bug fixes, a refactor, a dependency upgrade and one legacy-language task) and freeze them as your test set.
This Month:
- Run the 20 tasks three ways: SaaS Bob on a sanitized copy of the repository, self-hosted Bob on Laguna S 2.1, and self-hosted Bob on Nemotron if you have the hardware. Score pass rate, reviewer edits per task and wall-clock time. A single model swap can move results without warning, so pin the weights version you test.
- Have your SOC confirm it can capture Bob's activity from OpenShift logs, since Bob itself does not provide security event logging in the self-hosted build.
- List every workflow that relies on screenshots or diagrams. Those break on both local models.
Before Renewal or Signature:
- Write the pass-rate gap into the business case. If self-hosted Laguna closes 14 of 20 tasks and SaaS Bob closes 18, price that gap in reviewer hours.
- Negotiate a model-substitution clause: if IBM adds a stronger local model, your license covers it at no extra cost.
The Bottom Line
IBM has done what regulated buyers asked for and moved the agent inside the network. The model running there is a different model from the one behind SaaS Bob's published results, and nobody outside IBM has measured what that costs in output.
This is the familiar on-premises trade from every earlier software cycle: you get control of the data and you give up whatever the vendor's cloud service was doing for you behind the interface. Here that is IBM's multi-model routing and access to Claude. For many banks and agencies the trade is worth making. Make it with your own pass rate on your own tickets, and with IBM's price, GPU sizing and hybrid controls in writing.
Continue Reading
- Cursor Alternatives That Survive Security Review Start With Copilot
- Best Air-Gapped LLM Stack: Apache Weights on vLLM Beat NVIDIA's Fee
- Nvidia Hired 109 Poolside Engineers. No Clause Fired.
- Tencent Says 49B Active. Your Node Loads All 770B.
- Barclays Picks Claude for Half Its Developers, Copilot for Desks
- IBM Cut a 9-Month Project to 3 Days With AI Agents
