If you prototyped on the OpenAI Agents SDK and now want out, check first whether the SDK is actually what holds you. The library itself is MIT-licensed and advertises support for "100+ other LLMs". What ties you to OpenAI is three features you probably turned on by default: hosted tools, server-side conversation state, and traces that land in OpenAI's dashboard. Unwind those three first. If you still need to move after that, move Python agents to Pydantic AI with Temporal or DBOS underneath; pick LangGraph only if you genuinely need a checkpointed graph; pick Mastra if your team writes TypeScript. The option to avoid, if lock-in is your reason for leaving, is the Claude Agent SDK — it trades one vendor's runtime for a deeper one.
| Option | Models it runs | Durable execution | Where traces go by default | Do NOT pick it if… |
|---|---|---|---|---|
| OpenAI Agents SDK (de-risked) | OpenAI natively; others via adapters | Temporal (GA), DBOS, Restate, Dapr integrations | OpenAI Traces dashboard | Most of your traffic will run on non-OpenAI models |
| Pydantic AI | Multi-provider by design | Temporal, DBOS, Prefect, Restate, AWS Lambda | Wherever your OpenTelemetry goes | You need a visual graph and time-travel debugging |
| LangGraph | Multi-provider | Built-in checkpointers (Postgres, SQLite) | Wherever you send it; LangSmith is the path of least resistance | Your agent is a loop with tools, not a graph |
| Mastra | Multi-provider via provider/model routing |
Workflow suspend/resume snapshots in your storage | Your configured exporter | Your agent code and team are Python |
| Google ADK | Gemini first; others via adapters | Your own infra or Google's managed runtime | Your exporter, Google Cloud if you deploy there | You are not already on Google Cloud |
| Claude Agent SDK | Claude | Sessions you resume; no workflow engine | Your exporter | Your reason for leaving is vendor lock-in |
Pricing and capabilities as of September 30, 2026, from each vendor's documentation and pricing page, linked below.
Where Is the OpenAI Agents SDK Actually Provider-Specific?
The provider-specific parts are the Responses-only tools, the hosted tools and the OpenAI-backed session — not the Agent, Runner or handoff primitives. Those primitives are ordinary Python, and they are the part most teams assume they will have to rewrite.
Walk your codebase for these, in this order:
- Responses-only features. The SDK's own models documentation lists
ToolSearchTool,tool_namespace(), deferred-loading tool surfaces andProgrammaticToolCallingToolas features that "are rejected on Chat Completions models and on non-Responses backends." Any of them in your code is a hard dependency. - Hosted tools. Web search, file search and code interpreter run on OpenAI's side and bill there. As of September 30, 2026, OpenAI's pricing page lists web search at $10.00 per 1,000 calls plus search content tokens, file search at $2.50 per 1,000 calls plus $0.10 per GB per day of storage (1 GB free), and code-interpreter containers from $0.03 per 20-minute session at 1 GB. Leaving means replacing each with your own search API, vector store and sandbox — and re-indexing every file you uploaded.
- Server-side state. Of the session backends in the sessions documentation, only
OpenAIConversationsSessionstores history on OpenAI's servers; SQLite, Redis, SQLAlchemy, MongoDB and Dapr all keep it on infrastructure you control. If you picked the OpenAI one, your conversation history is an export job, not a config change. - Default tracing. Covered in its own section below, because it is the one teams forget.
A definition, since the rest of this page depends on it: an agent runtime is everything that keeps an agent alive between model calls — state, retries, resumption, tool sandboxes and telemetry. The SDK is the API you write against. The runtime is where the lock-in sits.
The strongest case for staying is real. Non-OpenAI models plug in through three built-in integration points — a global client, a per-run ModelProvider, or a per-agent model. If you mostly run OpenAI models and want an escape hatch, switching sessions to Postgres and adding your own trace processor gets you most of the way for a fraction of a rewrite.
The weak spot is multi-provider tool calling. The SDK's issue tracker has reports of Claude tool calls failing ID validation through LiteLLM and tool calling breaking with Claude thinking models via LiteLLM. The SDK docs themselves warn that providers without structured JSON output "will occasionally produce invalid JSON." If your plan is that most of your traffic runs on another vendor's model, you are fighting the SDK's design, and that is when to leave.
Which Alternative Handles Durable Execution and Resumability?
Pydantic AI has the widest choice of durable engines; LangGraph has the most built in; the OpenAI Agents SDK gets durability only from an external engine. Durable execution means that a run which crashes at step 9 of 12 resumes at step 9 — without re-paying for steps 1-8 in tokens or repeating their side effects.
- OpenAI Agents SDK + Temporal. Temporal declared the Python integration generally available on March 23, 2026; its integration documentation still marks streaming, sandbox support and OpenTelemetry as pre-GA. Its listed limits matter:
LocalShellToolandComputerToolare not supported,SQLiteSessioncannot hold workflow state once workers are distributed, and MCP servers run outside the workflow. The SDK's running-agents documentation also lists DBOS, Restate and Dapr integrations; Temporal's is the one with a GA announcement. - Pydantic AI. Its durable execution page lists five co-maintained backends — Temporal, DBOS, Prefect, Restate and AWS Lambda — plus external integrations for Airflow and Kitaru. It also says, correctly, that "durability is not storage": the engine survives a crash inside one run, but saving a conversation to resume next week is a separate job.
- LangGraph. Durability comes from checkpointers — Postgres for production, SQLite for development, in-memory for tests. The docs warn that in-memory checkpoints vanish on restart and that checkpoints accumulate over long conversations, raising latency and storage costs. Plan pruning on day one.
- Mastra. Workflows can
suspend()andresume(), with snapshots stored in your configured storage provider that survive restarts and deploys. That covers human approval steps well; it is not a replay engine like Temporal. - Claude Agent SDK. It gives you sessions you can resume or fork, which is conversation continuity, not crash-safe workflow execution.
The position: if a run can pause for a human for hours or days, put a real workflow engine underneath. Temporal and DBOS each work with both the OpenAI SDK and Pydantic AI, which means you can adopt it before you migrate — and the migration then becomes swapping the agent inside a workflow you already trust.
Does MCP Make Tools Portable Between SDKs?
Yes for tools you host yourself; no for the vendor's hosted tools. Every option on this page consumes MCP: the OpenAI SDK has built-in MCP server tool calling, Google ADK supports MCP alongside A2A, and the Claude Agent SDK lists MCP as a core capability.
So the cheapest migration insurance you can buy this quarter is structural: move every business tool behind an MCP server you operate, even while you stay on OpenAI. Function tools defined inline as Python decorators have to be re-registered in each framework's syntax. Tools behind MCP move with a config change. The tools that stay stuck are the hosted ones from the section above — search, file search and code execution — because they are not tools you own; they are services you rent. Before you add MCP servers, read our MCP governance guide on where to enforce the allowlist.
What Observability Do You Keep When You Switch?
You keep only what you exported to a backend you control. By default the OpenAI Agents SDK sends traces to OpenAI's backend, visible on the platform Traces dashboard. The same page notes that tracing "is unavailable for organizations that use OpenAI's APIs under a Zero Data Retention (ZDR) policy," and that set_trace_processors() replaces the default exporter entirely.
Do that now, whether or not you migrate. Route spans to a backend you own — the SDK lists processors for Langfuse, LangSmith, Braintrust, Datadog, MLflow, Pydantic Logfire and others — and the history you would otherwise leave behind comes with you. Our agent monitoring comparison covers which backend to pick.
Two costs to know, as of September 30, 2026:
- Pydantic Logfire: Team plan is $49 per month with 10 million records included and $2 per million after; Growth is $249 per month, per Pydantic's pricing page.
- LangSmith: Plus is $39 per seat per month with up to 10,000 base traces a month, then pay-as-you-go; base traces keep 14 days, extended traces 180 days, per LangChain's pricing page. The per-trace overage rate is not on that page — ask sales before you size it.
One caveat on "just use OpenTelemetry": the GenAI semantic conventions that define invoke_agent and execute_tool spans were, as of July 17, 2026, still in Development status, with nothing marked Stable. Attribute names can still move. Export raw spans and keep them; do not build dashboards you cannot rebuild.
Why the Claude Agent SDK Is the Wrong Exit
The Claude Agent SDK is a good product aimed at a different buyer, and for a team leaving over lock-in it is the losing choice. Anthropic's documentation describes it as "Claude Code as a library" — a library that runs the Claude Code binary, with its tools, permissions, hooks and sessions. Its use is governed by Anthropic's Commercial Terms of Service, and third-party developers may not offer claude.ai login or rate limits in their products without approval.
That is a heavier coupling than the one you are escaping. The OpenAI SDK is a thin MIT library you can point at other models; the Claude Agent SDK is a vendor agent harness you embed. If you are building a coding or file-manipulation agent and have already standardised on Claude, it is excellent at that job. If your board asked you to "reduce dependency on a single model vendor", it moves you sideways.
Google ADK is the runner-up loser for the same reason, but milder. ADK is open source, ships SDKs for Python, TypeScript, Go, Java and Kotlin, and runs other vendors' models through adapters. Its gravity is Google Cloud: deployment is smoothest to Google's managed agent runtime, Cloud Run or GKE, and the managed runtime is metered per vCPU-hour and GiB-hour on Vertex AI's pricing page. Already on GCP? It is a reasonable pick. Not on GCP? You are choosing your next lock-in on purpose.
Who should NOT pick each winner:
- Pydantic AI — not for teams whose agent is really a state machine with many branches. You will end up hand-building the graph engine LangGraph already ships. It also has a smaller ecosystem than LangChain's, a trade-off Speakeasy's framework comparison flags alongside its type-safety strengths.
- LangGraph — not for a single agent with tools and one approval step. You will pay the graph's conceptual tax for nothing.
- Mastra — not for Python teams. Rewriting the language and the framework at once doubles the risk.
What Does a Realistic Rewrite Cost?
For most teams the rewrite is a few weeks of engineering, and the expensive part is re-validating behaviour, not porting code. State the workload so this is like for like: four agents, two handoffs, eight tools (three already behind MCP), about 50,000 runs a month, with a fifth of runs pausing for human approval.
For that workload, the order of effort — our estimate, not a benchmark — runs:
- Cheapest — de-risk in place. Swap to a Postgres or Redis session, replace the trace exporter, move inline tools behind MCP. Days, not weeks. No behaviour change to re-test.
- Moderate — replace hosted tools. Stand up your own search API, vector store and sandbox; re-index uploaded files. Every retrieval-dependent answer needs re-evaluation, because a different index returns different chunks.
- Most expensive — port the orchestration. Handoffs become Pydantic AI delegation or LangGraph edges; guardrails need re-implementing; prompts tuned for one model need re-tuning for another. Pydantic's own migration guidance is the right method whichever framework you pick: trace a real request, port "the smallest complete path," and test old and new "at the same boundary."
Budget the eval set before the code. A migration without a frozen regression set is a model change you cannot measure — our LangChain alternatives guide makes the same argument for fixing durability before you rewrite anything.
How to Decide: The Criteria That Predict Regret
Pick on where state lives and where traces go, not on API ergonomics. Every framework here can call a model and a tool. The teams that regret a choice regret its runtime.
- Will most traffic run on non-OpenAI models within 12 months? Yes → leave for Pydantic AI (Python) or Mastra (TypeScript). No → de-risk in place.
- Do runs pause for humans or last longer than a request? Yes → Temporal or DBOS under Pydantic AI, or LangGraph with a Postgres checkpointer. Never in-memory.
- Is the agent a graph with branching, retries per node and time-travel debugging? Yes → LangGraph. Otherwise the graph is overhead.
- Does anything use hosted search, file search or code interpreter? Price the replacement before you commit; it is often the largest line item of the migration.
- Are you a ZDR customer? Default tracing is already unavailable to you — you need your own exporter regardless.
This Week: run grep for OpenAIConversationsSession, WebSearchTool, FileSearchTool, CodeInterpreterTool and ToolSearchTool, and list every hit. That list is your lock-in.
This Month: replace the trace processor and the session backend. Move inline tools behind MCP.
Before Q4 Close: freeze a regression eval set, then port one complete request path to the target framework and compare outputs at the same boundary.
The Bottom Line
The last cycle taught the same lesson: teams that "left" a cloud by rewriting application code discovered the lock-in lived in the managed queue, the proprietary database and the logging pipeline. Agent SDKs repeat it one layer up. The Agent class was never the problem. The state, the tools you rented and the traces you did not export are.
Move those three, and leaving becomes optional.
Continue Reading
- LangChain Alternatives: Fix Durability Before You Rewrite
- LangGraph vs CrewAI vs AutoGen: One of Them Is Retired
- Your Assistants API Dies Aug 26. Azure's Exit Is Different.
- Agent Orchestration Platforms: Score Exit, Not Features
- AI Observability Pricing: Same 10 GB, $49 or $930
- The 2026 Agentic AI Stack: 8 Layers, 3 You Can Skip
