Reliable voice agents come from splitting the call into nodes wherever the business needs a guarantee, such as exact wording or a tool that must run first. Inside each node, keep prompts short and let the AI talk. Exact words come only from fixed text or recorded clips. Then test each leg, meaning each step of the call, several times.
Key Takeaways
- Split the flow where a guarantee is needed. Let the AI talk inside nodes.
- Exact words come from greeting text, transition lines or recordings, never a prompt.
- A leg that carries a guarantee must pass on every repeat run.
This post is part of our guide to testing, monitoring and scaling AI voice agents. The guide covers the whole loop, from testing before launch to scaling after it. This one is about where the nodes in a flow should go. We build Dograh, an open-source voice agent builder where every call is a graph of nodes.
One big prompt or a flow
One prompt works well for a short call, and it starts to slip as the call gets longer and the rules pile up.
For a short, low-stakes call, such as giving opening hours or taking a message, one prompt is often the right choice. Splitting a call like that into steps mostly makes the agent sound stiff.
Long calls are where it breaks. In a 2025 study by Microsoft and Salesforce researchers, 15 language models across more than 200,000 simulated conversations did 39% worse on average when a task arrived over several turns instead of all at once. Their ability dropped only 16%. Unreliability rose 112%, and once a model took a wrong turn, it rarely found its way back.
Rules have the same problem. The IFScale benchmark gave models up to 500 instructions at once. The best models followed nearly all of them when there were only a few, and slipped further as the count climbed into the hundreds. It measured a writing task, not phone calls, so read it for the shape of the curve: every rule you add to one prompt competes with all the others.
Our rule: split where the business needs a guarantee, and let the AI talk freely inside each piece.
Split the call where the business needs a guarantee
A new node earns its place when something has to happen the same way on every call.
Four kinds of moments qualify. Some words must be said exactly, like a recording notice or a legal disclosure. Some tools must run before the call moves on, like an identity check before anyone discusses an account. Some calls must reach a person when the caller asks for one. And some actions can't be undone, like taking a payment, so they need a confirm step of their own.
Everything else can stay inside a node. Small talk and rephrasing a question the caller didn't follow are the AI's job, and it does them better with room to talk.
In Dograh, a flow is a graph. A Start Call node opens the call, Agent nodes handle each stage, End Call nodes close it, and one Global node holds a prompt the other nodes share. Nodes connect through pathways, and each pathway has a condition written in plain language. The AI reads the conversation and takes the pathway whose condition fits best. Each node has its own prompt and its own tools, so a node only carries what its step needs.
OpenAI's voice prompting guide points the same way: when a call moves to its next phase, swap in the prompt and tools that phase needs.
Where exact wording comes from
Inside a node, the AI writes every sentence, so a prompt can ask for exact wording and still get a paraphrase.
Prompts leak by paraphrase. Tell the model to read a disclosure word for word, and on most calls it will. On some, it will shorten the line or fold it into the next sentence. A ban on a phrase fails the same way. The model avoids the phrase and still says what it means.
In Dograh, exact wording comes from three places. The Start Call node's greeting text is a fixed line spoken the moment the call connects, before the AI's first turn, and the greeting can be a recording instead. A pathway can carry its own transition speech, an optional line the agent says as it moves to the next node. And pre-recorded clips can be added to a node's prompt, so the agent plays a real recording. There is no separate fixed-line node.

Each guarantees something different. The greeting plays on every call. Transition speech plays on every call that takes that pathway, which makes it a good home for a disclosure tied to a step, such as the line before a payment. A clip fixes the words, while the AI still decides when to play it, so check the call record to confirm it played. Text transition speech does not play on speech-to-speech models.
The opening line deserves the most care. In missed-payment recovery in BNPL, the first sentence can decide which collection rules apply, so it should never be left to the model. It also sets how the caller judges the agent, as we cover in the first 15 seconds of an outbound call. Recordings cut text-to-speech spend too, which we work through in mixing pre-recorded audio with TTS.
Open Source Alternative to Vapi / Retell
Self-hosted voice agent platform — no per-minute fees
dograh-hq/dograh
Star on GitHub
Keep each node's prompt short, with only the tools it needs
A short node prompt is easier for the model to follow, and the conversation history arrives with it anyway.
Every time Dograh calls the model, it sends the whole conversation so far, along with the Global prompt and the current node's prompt. You don't need to repeat what the caller said earlier. The node prompt only has to say what this step is for and when it is done.
Short matters because long inputs hurt. Chroma's Context Rot study tested 18 models in 2025 and found that focused prompts of about 300 tokens scored well above full prompts of about 113,000 tokens on the same memory questions. A phone call never gets that long, but the lesson carries over. Anything the step doesn't need is text the model has to read past.
Tools follow the same rule. A node that can only check availability can't book a slot by mistake.
Write the node prompt like a briefing for a new colleague, since a wall of NEVER and ALWAYS rules tends to make agents worse. A line like "Never transfer the call without permission" can leave an agent refusing a caller who plainly asks for a person. "Transfer the call when the caller asks for a person" carries the same rule and tells the model what to do. OpenAI's guide also warns about words like never and always. Order helps as well: "confirm the date, then check the slot" reads better to a model than a tree of if-then rules.
Pathway conditions need the same care. If the agent keeps taking the wrong pathway, make the condition more specific. "Caller confirms the new date" works far better than "caller is done".
Give every node a way out
Callers will ask things nobody planned for, so every node needs a path for the unexpected.
After launch, callers also talk to an agent differently than they would to a person, in ways that are hard to predict while building.
So plan the exits. Give each Agent node a way to transfer the call to a person, and a pathway to an End Call node that closes the call politely. Dograh already requires every path to end at an End Call node. Put the handling for common off-script moments, such as a caller asking whether they are talking to an AI, in the Global prompt, so every node shares it.
Join the Dograh Community
Dograh is an OSS alternative to Vapi. Join our Slack community for queries, releases, best practices & community interactions.
Test each leg several times, then expect to rebuild
A flow is only reliable once each leg passes again and again, and it will change once real callers arrive.
The same prompt and conversation can produce a different reply on the next run, as the main guide shows, so one passing test call proves little.
A leg that passes four runs out of five sounds fine. At real call volume, it could fail on about one call in five. For the legs that carry a guarantee, set a higher bar: a leg passes only when it passes on every one of several runs in a row. If it fails even once, tighten the condition or split the step.
Langfuse helps here. Connected to Dograh, it lets you rerun one AI step with the exact prompt it received, as many times as you like, without dialling again. Then check the path as well as the words. After a call, View Trace shows which pathway the AI took on each turn, so you can confirm the call went through the disclosure step. You can't prove a model will never say something, but you can prove which path every call took. Our guide to testing voice agents before go-live covers the full routine, including a test over a real phone line.

Then expect to rebuild. Teams often rework the whole flow, and sometimes the use case, once the agent goes live. Small nodes make that cheaper, because you replace one step instead of rewriting one long prompt. In Dograh, each change is saved as a draft while the published version keeps taking calls, and a separate test URL runs the latest draft. After any change, rerun every leg it touches.
This is also why the flow should be something your team can read. Dograh is the orchestration layer of a voice agent: it holds the flow, the turn-taking, the tool calls and the telephony, while the speech and language models are a separate layer you choose. It is open source under the BSD-2 licence and can run on your own servers, so the flow and every call's trace stay yours. When a regulator asks how a call was meant to go, you can show them the flow, and the trace shows the path each call actually took.
Start with the one step that must never go wrong. Give it its own node, put the exact words in transition speech or a recording, and build outward from there.
Glossary
- Pathway condition
- The plain-language test on a pathway between two nodes. The AI reads the conversation and takes the pathway whose condition fits best, so a vague condition sends calls the wrong way.
- Transition speech
- An optional fixed line attached to a pathway and spoken as the agent moves to the next node. Every call that takes that pathway hears the same words.
- Global prompt
- One prompt per workflow that Dograh adds to each node's own prompt when Add Global Prompt is on. Rules that hold for the whole call go here once.
- Greeting text
- The fixed line a Dograh agent speaks the moment a call connects, set on the Start Call node. It plays the same words on every call, so it is the place for a line that must be said exactly.

