create AI voice agent

From Zero to a Real AI Call in 3 Easy Steps with Zoice AI

SB

Shivansh Bhardwaj

AI Product & Systems Engineer

October 9, 202611 min read
From Zero to a Real AI Call in 3 Easy Steps with Zoice AI

You have work that happens on the phone: appointment confirmations, payment reminders, and inbound questions that arrive after your team has gone home. You want an AI to take those calls. Search for how to create an AI voice agent and you get two kinds of answers. One is an architecture diagram with five vendors wired together. The other is a demo video that never touches a real phone number. Neither tells you what has to happen between the idea and a customer's phone actually ringing.

That gap matters because a voice agent is judged on a live call, not in a test console. The caller does not care which speech model you picked. They notice whether the agent answered the number they dialled and understood them the first time. They notice whether it paused too long, and whether it could get them to a person when the question was beyond it. Each of those depends on a separate piece of plumbing. A missing piece is what turns a good demo into a bad first week in production.

This guide breaks the work into three steps: define the agent, connect a number, and test before you go live. First, it helps to see what sits underneath, because the decisions in step one only make sense once you can see the parts they configure. The examples use the Zoice AI phone calls platform, where all three steps are set up without code. The same structure applies to any conversational AI stack.

The Four Jobs Inside Every AI Voice Agent

Underneath the interface, every AI phone agent does four jobs: carry the call, work out what the caller said, decide what to say, and say it. Most of the setup work is choosing how each one is handled:

  • Telephony carries the call over SIP, the Session Initiation Protocol, which OpenAI's documentation describes as a protocol “used to make phone calls over the internet” and which is standardised in RFC 3261 (2002). A real number usually comes from a SIP trunking provider, which converts phone-network traffic into IP.
  • Listening turns the caller's audio into a transcript, usually through speech-to-text (STT).
  • Deciding is done by a large language model (LLM), guided by your instructions and ideally your own documents.
  • Speaking turns the reply back into audio through a text-to-speech (TTS) voice, which shapes much of the caller's impression of the call.

There are two common ways to connect the middle three jobs. OpenAI's voice agents guide (as of October 2026) describes a chained pipeline: speech-to-text, then the agent, then text-to-speech. It sums up the benefit as: “Inspect or transform intermediate text and replace each component independently.” The alternative is speech-to-speech, where a single model is used “to interpret audio, decide what to do, and respond in speech”. The same guide now lists a third option, GPT-Live, for full-duplex conversations with a separate backend.

ApproachHow it worksUsually suits
Chained pipeline (STT → LLM → TTS)Three stages joined by textTeams that need transcripts, redaction, and the freedom to swap a speech or voice vendor
Speech-to-speechOne model hears the audio and replies in audioCalls where expressive audio and fewer handoffs between stages matter more than control over each stage
Full-duplex with a separate backendBoth sides can speak at once; the backend runs separately from the voice modelDesigns that put natural overlapping talk first

The “usually suits” column is engineering guidance rather than benchmark data; the right choice depends on your languages, your compliance needs and how much delay your callers will tolerate. If you need to keep anything as text after the call, such as a transcript or a redacted record, make sure the design produces reliable text somewhere in the flow before you choose a voice.

Key Insight

Creating an AI voice agent comes down to three jobs. Write instructions you could check against a transcript, connect a real number so calls reach the agent, and test on an actual phone line before any customer hears it. The problems callers notice most, such as long pauses, poor recovery from interruptions and broken handovers, usually only show up on a real call. That makes the test step the one you cannot skip.

On this page

Step One: Define What the Agent Does

Most of a call's quality is decided in the agent definition, and that is mostly writing rather than engineering. On Zoice you set up the agent's script and behaviour in the dashboard. The same definition then runs on phone calls, WhatsApp and the web chat widget, so it is worth getting right once.

Write instructions like a job description

A vague prompt produces a vague agent. Treat the instructions as the brief you would give a new hire on their first shift. Make every line something you could check against a transcript afterwards. A workable brief covers:

  • The persona and the single job of the call. For example, “confirm tomorrow's appointment and offer one reschedule slot” rather than “help the customer”.
  • What the agent must collect. Zoice extracts structured fields you define from every call, so name each one for the system that will use it: for example, a preferred callback time for your CRM or an order reference for finance.
  • The objections you expect and the answer to each, plus the lines the agent must never cross, such as quoting a price it has not been given.
  • When to hand over. Zoice agents have a transfer_agent tool that passes the conversation to another AI specialist, and a transfer_number tool that transfers to a human phone number with an optional announcement.

Write objectives alongside the instructions: the outcomes a successful call must achieve. Every call then reports objective achievement results, so you can compare test calls against the same criteria live calls will face. That is what makes step three measurable rather than a matter of taste.

Choose a voice and language you have actually heard

Language and voice are the first things a caller judges. Zoice covers Hindi, Bengali, Punjabi, Telugu, Marathi, Tamil, Gujarati, Kannada, Malayalam, Odia and English (India). It also handles Hinglish natively, for callers who mix Hindi and English in the same sentence. Every catalog voice has an instant preview, including with up to 200 characters of your own text. Listen to your real opening line rather than a sample sentence. You can also pitch-shift a voice from -12 to +12 semitones and add light background ambience, such as office hum, so the line does not sound unnaturally silent.

If the agent will answer questions about your products or policies, attach a knowledge base instead of pasting everything into the prompt. Answers are then retrieved from your documents. The retrieval test bench shows the exact chunks the agent would receive for a query before any caller hears an answer built on them. This also makes the prompt shorter and easier to review, because the facts live in documents you already maintain.

Core Components

1Agent Instructions and Objectives

A written brief covering the persona, the single job of the call, the objections to expect and when to hand over. It also sets the objectives each call should achieve, and every conversation reports objective achievement results against them.

2Structured Extraction Fields

Structured fields you define, such as a callback time or an order reference. They are extracted from every call alongside the AI summary and sentiment verdict.

3Voice and Language

A catalog voice previewed with your own opening line, in Hindi, Tamil, Telugu, Bengali or another supported Indian language, or English (India). Pitch shift and background ambience are optional.

4Phone Number Connection

Either an imported Plivo number assigned to an agent, or your own SIP trunk over UDP, TCP or TLS with numbers bound to agents for inbound routing.

5Live Testing and Supervision

Per-call voice-quality scores, p95 turn latency and per-message latency logging to find slow turns. Whisper and voice takeover let a human step into early live calls.

Step Two: Connect a Real Phone Number

An agent with no number is a demo. Underneath, every AI phone agent connects in the same way: the phone network hands the call to a SIP endpoint, and the platform decides which agent answers. OpenAI's Realtime API with SIP guide shows the pattern for a build-it-yourself setup. You buy the number from a SIP trunking provider, and a webhook tells your application that a call has arrived. Your code then has to “approve the inbound call and configure the realtime session that will answer it”, including its instructions and voice. You build and run all of it yourself, from the webhook endpoint to handling the case where a session fails to start.

On Zoice, that routing is configuration rather than code. There are two routes, and the right one depends on whether you already have a contract with a phone carrier.

Import a number and assign an agent

The quickest path is to import a number and assign any AI agent to it. Inbound calls are then answered automatically, around the clock. Plivo numbers are supported today. Each imported number shows its webhook status, with a one-click retry if the connection needs to be set up again. It is the same webhook pattern as the build-it-yourself version, but managed for you.

Bring your own carrier over SIP

You may already have negotiated rates, or numbers your customers recognise. In that case, register your own trunk through SIP / BYOC calling. Trunks run over UDP, TCP or TLS, authenticate with digest credentials or an IP allowlist, and can encrypt the call audio with SRTP. Your numbers, known as DIDs (direct inward dialling numbers), are bound to agents for inbound routing. Every trunk has a test button that sends a real SIP OPTIONS probe to your carrier and reports the result. Your carrier keeps billing you directly at your negotiated rates, and Zoice bills for the AI conversation layer.

Outbound calls use the same agent. You can place a single call, upload a CSV for a bulk run, or set up a full campaign. In bulk uploads, each contact can be assigned their own language, which helps when one list covers callers in both Chennai and Lucknow. If you can, start with inbound calls. People who rang you themselves tend to be more forgiving test subjects than people you interrupted.

Implementation Roadmap

  1. 1Write the agent's instructions, objectives and extraction fields as a job brief, including the exact handover rules for transfer_agent and transfer_number
  2. 2Preview catalog voices with your real opening line, choose the language, and attach a knowledge base checked with the retrieval test bench
  3. 3Connect a number by importing a Plivo number or registering your own SIP trunk, then confirm the webhook status or OPTIONS probe result
  4. 4Test from an ordinary mobile phone: interrupt the agent, go silent, ask out-of-scope questions, trigger every transfer, and review each call's summary and objective results
  5. 5Go live with a small inbound or outbound batch, supervise with whisper and takeover, and widen the rollout only after quality scores and objectives hold

Step Three: Test the Call, Then Go Live

The last step is the easiest to rush, yet it decides whether the agent feels like a conversation or like a slow phone menu. Human turn-taking sets a high bar. Levinson and Torreira, writing in Frontiers in Psychology, note that gaps between turns are short, “of the order of 200 ms”. Producing speech takes much longer, “over 600 ms”, so people start planning their reply while the other person is still talking. Overlap is rare too: it occupies “less than 5% of the speech stream”.

Two practical lessons follow, and neither is to treat the human gap as a deadline. That figure is a description of people, not a target that today's agents usually hit. The first lesson is that every stage of the pipeline adds to a delay the caller can hear, so trace long pauses back to the stage that causes them. The second is that people rarely talk over each other. So when a caller does cut in, the agent must stop speaking promptly, and it must not mistake a thinking pause for the end of a turn. Detecting the end of a turn and handling interruptions are the techniques built for exactly these problems.

Before any real customer hears the agent, run a short, deliberate test pass:

  • Call the number from an ordinary mobile phone, not a browser tab. Listen to the opening line and the first response.
  • Interrupt the agent mid-sentence, then go silent for a few seconds, and check that it recovers both times.
  • Ask something the knowledge base does not cover and confirm the agent says so instead of making something up.
  • Trigger every handover path and confirm the announcement plays and the right person or agent picks up.
  • Read the AI summary, extracted fields and objective results for each test call, and fix the instructions wherever they are wrong.

Zoice gives each call measurements to check against. The analytics layer scores voice quality across latency, interruptions, silence and talk balance. It also reports p95 turn latency, a percentile view that surfaces the slow turns an average would hide, and logs tokens and latency per message, so you can find where a conversation slowed down. During the first live days, an operator can whisper instructions that only the AI hears. They can also trigger a voice takeover that dials their own phone into the call.

For outbound calls, compliance belongs in the test pass too. Zoice captures an explicit yes or no consent in the first 5 seconds of every outbound call, designed for India's Digital Personal Data Protection (DPDP) Act and the US Telephone Consumer Protection Act (TCPA) workflows. It can also redact patterns such as Aadhaar, PAN and OTP numbers before transcripts reach storage. The security page covers the rest. Then go live with a small batch, read the results, and widen the rollout only when the numbers hold.

Frequently Asked Questions

Can I create an AI voice agent without writing code?

Yes. On Zoice you configure the agent's script, connect a telephony number and go live from the dashboard, with no code required. Code becomes useful later, not at the start: integrations and webhooks are there when you want call outcomes to flow into a CRM or other systems.

Can an AI phone agent hand a call over to a human?

Yes. The transfer_number tool connects the caller to a human phone number, with an optional announcement such as “Connecting you to our support team”. A pre-transfer message briefs whoever picks up. Supervisors can also join live calls directly.

Is it legal to use an AI voice agent for outbound calls?

It depends on where you are and where the people you call are. Outbound calling is regulated in most markets: in India by Telecom Regulatory Authority of India (TRAI) rules and the DPDP Act, in the US by the Telephone Consumer Protection Act (TCPA). This guide is not legal advice, so have a lawyer review your script and consent flow before launch. Zoice's consent capture, recording redaction and TRAI-DLT-aware outbound campaigns are built for those workflows, but they do not replace that review.

Can one agent also answer WhatsApp and web chat?

Yes. The same agent configuration runs on phone calls, on WhatsApp through Meta's Cloud API, and in an embeddable web chat widget, and one knowledge base serves all three. Web chat conversations land in the same unified inbox as voice and WhatsApp, so your team supervises every channel in one place.

To see the three steps applied to your own calls, talk to the team and bring the call flow you want to automate first.

Sources

  1. Timing in turn-taking and its implications for processing models of language — Frontiers in Psychology
  2. Voice agents — OpenAI
  3. Realtime API with SIP — OpenAI
  4. RFC 3261: SIP: Session Initiation Protocol — IETF / RFC Editor
SB

Written by

Shivansh Bhardwaj

AI Product & Systems Engineer

Shivansh Bhardwaj is an AI Product & Systems Engineer at Zoice, focused on building and scaling production-grade AI products and conversational systems. His work spans AI agent architecture, real-time voice pipelines, LLM orchestration, latency optimisation, and multilingual AI experiences. He writes about emerging AI technologies, engineering challenges, and practical approaches to building fast, reliable, and scalable AI products for real-world business use cases.

AI ProductsConversational AIAI AgentsLLM EngineeringVoice AI

Ready to put an AI agent to work?

Deploy voice, WhatsApp, and chat agents across Indian languages — grounded in your knowledge and measured on every call.