From OpenAPI spec to agent tool, one factory across every tenant
An OpenAPI spec already describes an endpoint fully. A factory turns it into a callable agent tool, reusable across every tenant's backend.
- A tool with a backend's address and auth baked into its code only works for one tenant.
- Split config into a spec file (the API plus its base URL and auth) and a mapping file (which operations become tools, and how) per tenant.
- Hide server-fixed parameters from the model's schema and reshape flat model arguments into the real request body with a mapping file, no new code.
- A declarative flag, not a hardcoded tool-name list, is what lets a tool safely reuse one stateful resource (a cart, a draft) across a whole conversation.
- Load tools one at a time so one broken mapping costs one tool, not the whole agent.
Say an agent needs to call a backend: look up an order, cancel it, check stock. The fastest way to give it that ability is to hand-write one tool function per API endpoint. That works for the first backend. It breaks the moment a second tenant needs the same agent talking to a different backend, with a different address, different authentication, and slightly different rules.
Why one tool per endpoint stops scaling
A hand-written tool usually has the backend's address, the auth type, and a site or account id baked into the function body. Add a second tenant and you either duplicate the function with new constants, or add branching logic inside it. Both mean every new tenant is a code change, a review, and a deploy, even though nothing about the actual capability changed.
Generate the tool from the spec instead
An OpenAPI spec already describes an endpoint fully: the path, the method, the parameters, and their types. That is most of a tool definition already. The missing piece is not more information about the endpoint. It is a small amount of configuration saying which endpoints to expose, how their arguments map to what the model sees, and where each tenant's copy of the backend actually lives.
Split that configuration into two files per tenant: a spec file (the OpenAPI document plus its base URL and auth type) and a mapping file (a short list of which operations become tools, and how). A loader reads both at startup and hands them to a factory, which builds one callable tool per mapped operation. A new tenant only adds its own copy of these two files, the tool code itself never changes.
Each tenant has its own copy: the OpenAPI document itself, plus the base URL and auth type for that tenant's backend.
A short config file listing which spec operations become tools, with any fixed values, extra arguments, and body reshaping each one needs.
At startup, the loader reads every tenant's spec and mapping files and hands each mapped operation to the factory.
For each mapped operation, the factory builds a callable tool: its schema, its fixed values, and how to reshape a call into a real request.
The finished tool is registered against the agent it belongs to, ready to be called the moment a conversation needs it.
What happens once, at startup, before any conversation starts. Click a step for detail.
What the model sees versus what the request needs
A raw OpenAPI operation is not a good tool schema on its own. Some of its parameters are fixed per tenant, a site id, a store code, and must never be something the model decides. Others need renaming or nesting to match what the underlying request body actually expects. The mapping file handles both without touching the spec or writing a new function.
- Fixed values: filled in by the server, hidden from the model so it can never override them
- Extra arguments: not in the spec, but needed for the tool to make sense to the model (an order number, a quantity)
- Body mapping: reshapes the flat arguments the model gives into the nested structure the real request body expects
tools:
- tool_name: cancel_order
operation_id: doCancelOrder
description: "Cancel an order by its order number."
# filled in server-side, never shown to the model
param_defaults:
site_id: main-store
# arguments the model actually supplies
extra_params:
- name: order_number
type: string
required: true
# model's flat arg -> the request body's real field name
body_mapping:
code: order_numberThe model only ever sees the arguments the mapping exposed to it, nothing the server fills in itself.
Values like a tenant's site id are injected here, from config, never from anything the model supplied.
The model's flat arguments are reshaped into whatever nested structure the real request body expects.
The auth resolver adds the right headers for this tenant, whether that is a fixed token or a freshly resolved OAuth one.
The finished request goes to this tenant's own base URL, never another tenant's.
The backend's response comes back to the model as the tool's result, ready to reason over.
What happens on every single call, using the config built at startup. Click a step for detail.
Keeping one object stable across a conversation
Some tools are not one-off calls. They act on a piece of state that has to stay the same across a conversation: a cart, a draft, a session. If two calls create two different carts, viewing and adding no longer agree on which cart is which. Checking this by tool name, 'if the tool is called add_to_cart, resolve a cart first', works for the first backend and breaks for any tenant whose operations are named differently.
A steadier fix is a flag in the mapping itself: this tool operates on a bound resource. At load time, the factory resolves which operation creates that resource and which one lists existing ones, both from config too. At call time, it fetches or creates that resource once, caches the id per user, and reuses it for every bound tool. If a cached id turns out to be stale, the resource was removed or finished elsewhere, it clears the cache and resolves again, once, under a lock so two calls arriving at the same time cannot create two resources.
The fix was not a smarter cart lookup. It was making 'this tool needs one persistent id' a declared property of the tool, not a fact buried in its name.
Keeping tenants from leaking into each other
Two tenants can share the same auth client id on different backends, so a token cache keyed on client id alone can hand tenant A's token to tenant B by accident. Key it on tenant plus client plus token endpoint instead. And if a tenant's own credentials cannot be resolved, fail with an error rather than quietly falling back to a shared or default identity. A missing config should be loud, not a wrong-tenant response that looks like success.
One broken mapping should not take down the rest
Load tools one operation at a time, not all at once as a single unit. If one mapped tool points at an operation that does not exist in the spec, skip that tool and log why. Every other tool for that tenant, and every other tenant, still loads. A typo in one config file should cost one tool, not the whole agent.
The cost is a layer between the spec and the running tool that a single backend with a handful of endpoints does not need. Plain hand-written functions are simpler there, and the config buys nothing. It earns its cost once the same capability has to run against more than one backend or more than one tenant's account. From that point on, adding a tenant is a config change instead of a deploy.