The problem with deploying agents one site at a time
The original plan was simple: deploy a conversational lead agent on each client website individually. Each deployment would have its own configuration, its own prompts, its own CRM integration. This is how most people do it. It is also how most people end up with a fragile, expensive-to-maintain patchwork of separate deployments with no shared logic, no shared learning, and no real visibility into what any of them are doing.
After deploying two agents this way, the problems were already obvious. A change to the core conversation flow had to be applied twice. A bug in the lead qualification logic had to be found and fixed twice. When one client changed their CRM, that integration had to be rebuilt from scratch. The agents were not sharing anything — not infrastructure, not context, not configuration patterns. They were just two separate systems wearing the same costume.
The answer was to stop deploying agents and start deploying one agent platform that serves multiple deployments from a single orchestration layer — what became the Lead Agent Control Plane.
Why centralized instead of distributed
The alternative to a centralized control plane is a federation model: each deployment runs its own independent agent, and coordination happens at the configuration layer — shared prompt templates, shared playbooks, but separate execution environments. Some platforms do this. It has real advantages: blast radius is contained (a bad deploy affects one client, not all), and you can customize the execution environment per deployment without cross-contamination.
We rejected the federation model for three reasons. First, observability: when something goes wrong with a federated agent, you have to look in multiple places to understand what happened. A centralized plane means one log stream, one error surface, one place to look. Second, iteration speed: improving the conversation flow once and having it deploy to all active agents immediately is a significant operational advantage when you are still learning what works. Third, context sharing: a centralized plane can learn from conversations across deployments in a way that isolated agents cannot.
The question was not which architecture was technically superior. The question was which architecture we could actually operate well at this stage of the product.
The right answer depends on your operational maturity and the degree of isolation your clients require. For the Lead Agent Control Plane, centralized was the right call because we are building for small and medium businesses in Ghana and West Africa that do not need contractual data isolation — they need something that works reliably and improves over time.
How the control plane is structured
The control plane
The control plane has four main components: a deployment registry, a conversation router, a context engine, and an integration layer. Each one has a narrow responsibility and a defined interface to the others.
The deployment registry holds the configuration for each active agent deployment: which client, what persona, what qualification criteria, what CRM endpoint, what escalation rules. When a conversation comes in, the router queries the registry to pull the right configuration before doing anything else.
// Route an incoming conversation to the correct agent deployment
async function routeConversation(
request: ConversationRequest,
): Promise<AgentResponse> {
// 1. Identify deployment from request origin
const deployment = await registry.resolve(request.origin);
if (!deployment) {
throw new Error(`No deployment for ${request.origin}`);
}
// 2. Hydrate conversation context (session + visitor in parallel)
const context = await contextEngine.hydrate({
sessionId: request.sessionId,
deploymentId: deployment.id,
history: request.history,
});
// 3. Build agent prompt with deployment persona + context
const prompt = buildPrompt({
persona: deployment.persona,
context,
message: request.message,
rules: deployment.qualificationRules,
});
// 4. Run inference and return structured response
return agent.complete(prompt, { deployment });
}Routing logic
The router identifies a deployment from the request's origin — either a domain, an API key, or a widget embed token. The registry maps these to a deployment record. This sounds simple, and mostly it is. The complexity is in handling ambiguity: a client might have multiple domains (staging, production, a white-label version) that should all route to the same deployment, or they might have separate deployments per domain that share the same API key. The registry stores a many-to-one mapping from identifiers to deployments.
Context management
Context management is where centralization provides its clearest advantage. Because every conversation runs through the same context engine, we can track state across multiple sessions for the same visitor, carry qualification signals forward from previous conversations, and apply global rate limiting at the engine level rather than per deployment.
The context engine maintains two types of state: session context (short-lived, expires after inactivity) and visitor context (longer-lived, tied to a visitor identifier across sessions). Session context holds the conversation history and intermediate qualification signals. Visitor context holds persisted facts — job title, company size, use case — that have been confirmed in previous conversations and should not be re-elicited.
The worst experience in a sales conversation is being asked the same question twice. The context engine exists to prevent exactly that.
The stack
The control plane runs on Next.js with a separate Node.js service for the agent execution layer. The separation matters because the agent execution service runs long-lived inference requests and needs different scaling behavior from the Next.js API routes that handle the widget and dashboard UI.
# Lead Agent Control Plane — service topology
control-plane: # Next.js — dashboard, widget, API routes
runtime: node-20
deploy: vercel
agent-execution: # Node.js — inference, streaming, context IO
runtime: node-20
deploy: railway
scale: horizontal
context-store:
session: redis (upstash)
visitor: postgres (supabase)
registry:
primary: postgres
cache: redis (TTL: 5m)
integrations:
crm: per-deployment webhook
model: anthropic claude-haiku-4
fallback: claude-sonnet-4Claude Haiku handles the majority of conversations — it is fast enough for real-time interaction and cheap enough to run at volume. Sonnet is reserved for complex qualification conversations where a visitor is showing high intent signals and we can justify the latency and cost of a stronger model.
What broke in staging
Three things broke that we did not anticipate. All three are now fixed. All three are worth documenting because they are the kind of failures that do not show up until you run a real conversation against real infrastructure.
Context hydration latency. The context engine was making two sequential database calls — one for session context, one for visitor context — before the agent could start generating a response. In a streaming response model, this added 400–600ms before the first token appeared. That is a noticeable pause in a conversational interface. We moved both calls into a parallel Promise.all and brought the P95 hydration latency down to under 120ms.
Deployment registry cache invalidation. When a client updated their qualification rules from the dashboard, the router would continue using the cached config until TTL expiry. In a test where we changed the qualification criteria during an active conversation, the agent continued applying the old rules. We added a cache-bust event emitted by the registry on any deployment update, which the router subscribes to via a Redis pub/sub channel.
Persona bleed between deployments. This was the most unexpected failure. In our load tests, we sent concurrent requests from two different deployments to the same agent execution instance. Because the prompt builder was pulling deployment configuration from a module-level variable rather than scoping it to the request, high concurrency caused persona data from one deployment to occasionally appear in another deployment's prompts. The fix was obvious in retrospect: configuration must be scoped to the request context, not the module. Every prompt now carries its full configuration payload derived within the request boundary.
What the build taught us
Building the control plane changed how we think about agent deployment architecture in general. A few things that are now core to how we approach this:
The configuration layer is the product. The agent's conversational quality is table stakes — every serious AI platform has reasonable out-of-the-box conversation quality. The real differentiation is in how easily a non-technical operator can configure and tune the agent's behavior from a dashboard. The control plane UI is as important as the routing logic.
Streaming is not optional. Users tolerate a 200ms wait for a search result. They do not tolerate a 2-second wait before a conversational response begins. Streaming has to be designed in from the start — retrofitting it onto a non-streaming architecture is painful.
Qualification signals should be explicit, not inferred. Early versions of the agent tried to infer qualification from conversation tone and content — signals like "this person sounds senior" or "this message implies urgency." This is both unreliable and opaque. The current system requires explicit qualification criteria defined per deployment: specific questions the agent must ask and specific answers that trigger qualification. Explicit criteria are auditable, tuneable, and don't hallucinate.
What's next for the control plane
The control plane is near launch. What comes after launch is the evolution from a lead qualification system to a broader agent orchestration layer — one that can handle not just lead conversations but support queries, onboarding workflows, and internal business operations.
The architecture is already designed to support this. The deployment registry, context engine, and integration layer are all generic enough to power non-lead agent types. The first extension will be a support agent deployment using the same control plane infrastructure — different persona, different qualification logic, same operational foundation.
This note will be updated as the product ships and the production data starts coming in. There is always a gap between how you thought the architecture would behave and how it actually behaves when real users are in the conversation.