Evalgent
Back to Blog
Voice AI Evaluation

Grok Voice Agent Builder (xAI) Pricing: $0.08/min and Setup (Oct 2026)

Deepesh Jayal
Updated
17 min read
Grok Voice Agent Builder (xAI) Pricing: $0.08/min and Setup (Oct 2026)
On this page

Searchers want three things about the xAI voice agent builder: what it is, what it costs per minute, and how to set it up. This guide answers all three with numbers from xAI's own docs. It also covers what the price leaves out, where the agent can fail, and how to test it.

Every xAI price here was re-checked on October 5, 2026, against the xAI pricing page, which xAI last updated on September 29, 2026. Check it before you budget, because beta pricing can change. Our earlier version of this post quoted $0.05 per minute. xAI now lists $0.08 per minute for the current model, so we corrected it.

Evalgent is an independent evaluation platform. We do not resell xAI, and we cover it the same way we cover every provider.

What is the Grok voice agent builder?

Grok Voice Agent Builder: a no-code tool in the xAI console that turns a plain-language call description into a live voice agent. It runs on xAI's speech-to-speech model, Grok Voice.

xAI announced the builder in beta on July 1, 2026, per the launch post. The pitch is simple. You describe how calls should go, attach documents and tools, and pick a voice. xAI says you can reach a working agent in about two minutes.

The builder bundles the parts most teams assemble by hand:

  • Instructions: a plain-language playbook for greeting, resolving and wrapping up.
  • Knowledge base: documents grouped into collections the agent searches during calls.
  • Tools and connectors: your APIs, MCP servers, and apps like Google Calendar, Outlook, Linear, Notion and Google Drive.
  • Voice: built-in voices such as Ara, Eve and Leo, or a clone from about two minutes of audio.
  • Telephony: a free phone number per account, or your own number over direct SIP.
  • Guardrails: rules on what the agent must not say or do.
  • Call review: every call recorded and transcribed, with tool calls visible.

For compliance, xAI's Voice page lists SOC 2, HIPAA eligibility and GDPR compliance. The voice docs add that a BAA is available for healthcare use.

What model powers a Grok voice agent?

Under the builder sits the Grok Voice speech-to-speech API. The current flagship is `grok-voice-think-fast-2.0`, released July 29, 2026. The alias `grok-voice-latest` has pointed to it since August 5, 2026, per the Think Fast 2.0 announcement.

Speech-to-speech model: a single model that takes caller audio in and returns agent audio out. It replaces the usual chain of speech-to-text, an LLM, and text-to-speech.

That design removes two hand-offs per turn. It is the same class of system we compare in cascading vs speech-to-speech voice agents. The model also reasons while it speaks. Reasoning is on by default, with `reasoning.effort` set to `"high"`.

xAI publishes these vendor-reported figures for Think Fast 2.0, sourced to Artificial Analysis:

BenchmarkGrok Voice Think Fast 2.0Think Fast 1.0GPT-Realtime-2.1 (High)Gemini 3.1 Flash (High)
AA Speech-to-Speech Quality Index82.9%75.7%79.1%69.5%
Full Duplex Bench95.1%77.8%95.7%74.3%
Ï„-voice Bench (agentic)56.5%52.1%45.7%37.7%
Time to first audio0.70s1.25sNot listed2.98s

Treat these as a starting signal. Benchmarks use shared scenarios, not your callers, your tools, or your phone lines.

$0.08
Grok Voice audio rate per minute (xAI, as of Oct 2026)
$0.01
Extra per minute on a free xAI phone number
0.70s
Time to first audio for Think Fast 2.0 (vendor-reported)
~2 min
Time to a first working agent, per xAI

Grok voice agent builder pricing: cost per minute

Pricing as of October 5, 2026. Verified on xAI's pricing page (last updated September 29, 2026) and the builder launch post. Check current pricing before you commit.

ItemPriceSource
Speech-to-speech audio (`grok-voice-think-fast-2.0`)$0.08 per minute ($4.80 per hour)xAI pricing
Text input into a voice session$0.004 per text inputxAI pricing
Telephony on a free xAI number$0.01 per minuteBuilder launch post
Voices, including built-in voicesIncludedBuilder launch post
Platform feeNoneBuilder launch post
Collections search (RAG) tool$2.50 per 1,000 callsxAI pricing
Web search tool$5 per 1,000 callsxAI pricing
Collection storage$0.10 per GiB per dayxAI pricing

xAI states that builder agents bill "at our API rate." So the builder and the raw API share one audio meter. The builder adds no separate platform charge.

Grok voice models and pricing

xAI's Voice API has three metered modes. The builder uses the speech-to-speech model. Prices are as of October 5, 2026, from the xAI models page.

Grok voice model or modeWhat it doesPrice
`grok-voice-think-fast-2.0` (alias `grok-voice-latest`)Speech-to-speech voice agents, used by the builder$0.08 per minute ($4.80 per hour), plus $0.004 per text input
Speech to Text (default `grok-voice-transcribe-2.0`)Batch and streaming transcription$0.10 per hour (REST), $0.20 per hour (streaming)
Text to SpeechText to spoken audio with built-in or cloned voices$15.00 per 1M characters

Grok Voice Think Fast 2.0 pricing is the only rate that matters for a builder agent. Speech to Text and Text to Speech bill separately only if you call those APIs yourself.

Grok voice agent API pricing vs builder pricing

Grok voice agent API pricing and builder pricing are the same audio rate. The difference is what you run yourself. With the builder, xAI hosts the playbook, knowledge base, phone number and call logs. With the raw Voice API, you host the client, the tool servers, and your SIP bridge. Your servers cost money, but xAI's bill stays at the same per-minute rate.

That makes Grok voice agent builder pricing 2026 budgets easy to model. Multiply minutes by $0.08, add telephony, then add any metered tools. Any xAI phone agent, built in the console or in code, follows that formula.

Two items need care. First, the tool and storage rates come from the general API pricing page. xAI does not say whether builder calls meter each knowledge search this way. Check your console usage after a test batch. Second, a pure speech-to-speech model has no separate LLM line. That is a real difference from cascaded platforms, where the LLM is billed on its own.

Illustrative cost per minute

Price per minute is easy. Cost per call is what your finance team sees. Here is illustrative math for a support line with four-minute calls. All numbers below except the xAI rates are assumptions.

Line itemPer minutePer 4-minute callPer 1,000 minutes
Grok Voice audio$0.080$0.320$80.00
Free xAI number$0.010$0.040$10.00
Two knowledge searches per call (if metered)$0.00125$0.005$1.25
Illustrative total~$0.091~$0.365~$91.25

At 100,000 minutes a month, that is roughly $9,125 in xAI charges. That figure is illustrative. It excludes your own API hosting, escalation staff, and any carrier fees if you bring your own number.

What a Grok voice agent costs per minute: xAI model or builder charges, telephony, and tool or LLM costs, shown as an illustrative per-minute and per-1,000-minute breakdown

What changes the bill

A few settings move cost more than you expect:

  • Bring-your-own number. Over SIP, you pay your carrier instead of the $0.01 xAI rate. Twilio, Telnyx and Plivo each price minutes differently.
  • Silence and hold time. Billing is per minute of audio. Long holds during slow tool calls still count.
  • Idle re-engagement. The `idle_timeout_ms` setting makes the agent re-prompt silent callers. That keeps dead calls alive longer.
  • Transfers. A warm transfer keeps the session open until the destination answers.
  • Text turns. Injected text inputs bill at $0.004 each, so heavy scripted injections add up.

For a broader view of these line items, see voice agent pricing hidden costs and our AI voice agent cost guide.

From cost per minute to cost per resolved call

Cost per minute hides the metric that matters. What you pay for is a resolved call. Use this illustrative formula:

Cost per resolved call: total voice spend divided by calls the agent fully resolved without a human.

Take 1,000 four-minute calls at about $0.365 each. That is about $365. If the agent resolves 70% of them, you pay about $0.52 per resolution. At 50%, it rises to about $0.73. The per-minute price did not change. Quality did.

xAI reports a 70% autonomous resolution rate for Starlink support in its Think Fast 1.0 post. That is one vendor-reported deployment with 28 tools. Your rate will differ. Our guide to voice agent cost per resolution walks through the full model.

# Illustrative cost-per-resolved-call calculator (simplified)
AUDIO_PER_MIN = 0.08      # xAI rate, as of Oct 2026
PHONE_PER_MIN = 0.01      # free xAI number
SEARCH_PER_CALL = 0.0025  # collections search, if metered

def cost_per_resolved(calls, avg_min, searches, resolution_rate):
    per_call = avg_min * (AUDIO_PER_MIN + PHONE_PER_MIN) + searches * SEARCH_PER_CALL
    total = calls * per_call
    return round(total / (calls * resolution_rate), 3)

print(cost_per_resolved(1000, 4.0, 2, 0.70))  # ~0.521
print(cost_per_resolved(1000, 4.0, 2, 0.50))  # ~0.73

How to set up a Grok voice agent to answer phone calls

The builder flow is short, and it ends with a live phone number. These steps follow xAI's launch post and voice page. Button labels may shift during the beta.

How to build a Grok voice agent with xAI: configure the agent, voice and instructions, add tools, connect telephony or an app, then test and deploy

1. Open the builder. Sign in to the xAI console and open Voice Agents. Pick a preset such as Customer Support, Sales Associate, Personal Assistant, or Custom.

2. Write the playbook. Describe greeting, resolution and wrap-up in plain language. Keep it short. xAI's migration notes say Think Fast 2.0 needs simpler prompts than older models.

3. Attach knowledge. Upload policies and FAQs as documents. The builder accepts text, Markdown, Word, PowerPoint, Excel, HTML, JSON and PDF. Group them into collections you can share across agents.

4. Add tools and connectors. Wire your APIs, MCP servers, and apps like Google Calendar or Linear. Add a transfer-to-human tool and an end-call path.

5. Pick a voice. Choose a built-in voice or clone your brand voice from about two minutes of audio.

6. Set guardrails. List what the agent must refuse, such as reading back card numbers.

7. Connect a phone number. Attach the free xAI number included with each account, or route your existing number over direct SIP. This is the step that lets the agent answer phone calls.

8. Preview in the browser. Talk to the agent, change the playbook, and hear the change at once.

9. Test before launch. Run scripted and adversarial calls at volume, not just a few demo calls.

10. Launch and review. Go live, then replay recordings and transcripts to catch failures early.

That is the full no-code path from an empty console to a Grok voice agent answering real phone calls. Step 9 is where most teams cut corners.

Under the hood: the API behind the builder

Developers can skip the builder and call the same model directly. The speech-to-speech endpoint is a WebSocket at `wss://api.x.ai/v1/realtime`. It follows the OpenAI Realtime API event shape. xAI says most OpenAI Realtime clients work by changing the base URL.

Here is a simplified session setup. It mirrors what the builder configures for you.

# Simplified: configure a Grok voice agent session over WebSocket
import asyncio, json, os, websockets

async def run():
    url = "wss://api.x.ai/v1/realtime?model=grok-voice-think-fast-2.0"
    headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
    async with websockets.connect(url, additional_headers=headers) as ws:
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {
                "voice": "eve",
                "instructions": "You are the support agent for Acme. Verify the order, then resolve.",
                "turn_detection": {
                    "type": "server_vad",
                    "threshold": 0.85,
                    "silence_duration_ms": 600,
                    "prefix_padding_ms": 333
                },
                "audio": {
                    "input": {"format": {"type": "audio/pcmu"}},
                    "output": {"format": {"type": "audio/pcmu"}}
                },
                "tools": [{"type": "function", "name": "lookup_order",
                           "description": "Get order status by order number",
                           "parameters": {"type": "object",
                                          "properties": {"order_id": {"type": "string"}},
                                          "required": ["order_id"]}}]
            }
        }))
        async for msg in ws:
            print(json.loads(msg)["type"])

asyncio.run(run())

Pin a versioned model name in production. xAI recommends this, because the grok-voice-latest alias moves when a new model ships. The `silence_duration_ms` value above is illustrative. Tune it on your own calls.

Telephony over SIP

For phone calls through your own carrier, xAI's SIP guide documents three steps. Register a Direct SIP number, handle the signed `realtime.call.incoming` webhook, and join the call over WebSocket with its `call_id`. The guide includes setup steps for Twilio, Telnyx and Plivo.

The incoming-call webhook looks like this (trimmed):

{
  "type": "realtime.call.incoming",
  "data": {
    "call_id": "00000000-0000-0000-0000-000000000000",
    "sip_headers": [
      { "name": "From", "value": "+14155550100" },
      { "name": "To", "value": "+18005550199" }
    ]
  }
}

Call control uses REST. A transfer is a `refer` request:

curl -X POST "https://api.x.ai/v1/realtime/calls/$CALL_ID/refer" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"target_uri": "sip:agent@example.com"}'

A few documented details matter for testing:

  • Provisioning xAI phone numbers through the API is not supported. The free number comes from the console.
  • A `200` from `refer` means the destination answered. A `502` returns the carrier's SIP code, and the caller stays on the line.
  • Keypad digits (DTMF) reach the model as text on SIP calls. They flush on `#`, after 2.5 seconds idle, or when the caller speaks.
  • Session resumption lets a later call continue a conversation. History expires after 30 minutes of inactivity.

For trade-offs between phone and browser transport, see SIP vs WebRTC for voice agents.

Voices, languages, latency and turn-taking

Which voices does Grok offer?

xAI's Voice API page lists 28 built-in voices as of October 2026: Ara, Eve, Leo, Rex, Sal, Altair, Atlas, Aurora, Carina, Castor, Celeste, Cosmo, Helios, Helix, Iris, Kepler, Liora, Lumen, Luna, Lux, Naksh, Orion, Perseus, Rigel, Sirius, Ursa, Zagan and Zenith. Eve is the default.

Per the Text to Speech docs, every built-in voice can speak every supported language, and voice IDs are case-insensitive. The same voices work in speech-to-speech and text-to-speech. You can list them with `GET /v1/tts/voices`. Voices are included in the builder's $0.08 per minute rate.

Can you clone a voice in Grok Voice?

Yes. Per xAI's Custom Voices docs, you clone a voice from a reference clip of up to 120 seconds. xAI recommends 90 seconds or more, recorded in a quiet room with one speaker.

  • Limit: up to 30 custom voices per team, free to create in the console.
  • Where it works: a custom `voice_id` works anywhere a built-in voice does, including builder agents and the realtime API.
  • API access: creating voices through `POST /v1/custom-voices` is limited to Enterprise plans.
  • Availability: custom voices are currently available only in the United States, excluding Illinois.
  • Verification: xAI's Voice API page describes a passphrase plus speaker-embedding match before a clone is created.

Custom voices are scoped to your team and are not shared with other users.

Languages, latency and turn-taking

Languages. Marketing pages say 25+ languages. The speech-to-speech docs list 20+ with a named table, including English, Spanish (Mexico and Spain), Hindi, Arabic variants and Japanese. The model detects the caller's language automatically. You can bias transcription with a `language_hint`.

Latency. xAI reports 0.70 seconds time to first audio for Think Fast 2.0. That is a lab figure. Your real latency adds carrier legs, codec work, and tool round-trips. A slow order lookup can double the gap before the agent speaks.

Turn-taking. Turn detection uses server voice activity detection, set as server_vad, which also handles barge-in. You control the threshold, silence duration, and prefix padding. Set silence too short and the agent cuts off callers who pause mid-address. Set it too long and the agent feels slow. See turn-taking evaluation for how to measure this.

Tools during speech. xAI's docs warn about overlapping audio after tool calls. If your client requests the next response too soon, two answers can play over each other. Test this path on every tool.

Grok voice agent vs alternatives

The builder competes with other bundled and self-built stacks. This table compares billing structure, not quality. All prices are as of September 2026. Check each vendor's page before deciding.

Grok Voice Agent BuilderElevenLabs AgentsOpenAI Realtime APISelf-built cascade (LiveKit or Pipecat)
ArchitectureSingle speech-to-speech modelCascade: STT, your LLM, TTSSpeech-to-speech modelYour chosen STT, LLM and TTS
Main meter$0.08 per audio minute$0.08 per minute hosting on plansToken-based audio pricingEach vendor bills separately
LLM costIncluded in audio rateBilled separately by usageIncluded in token pricingBilled separately
TelephonyFree number at $0.01/min, or SIPProvider billed at costYou bring itYou bring it
LLM choiceGrok Voice onlyMany LLMsOpenAI models onlyAny
No-code builderYes (beta)YesNoNo

Sources: xAI pricing, ElevenLabs Agents pricing, and OpenAI pricing.

On grok voice vs elevenlabs, the key point is simple. The $0.08 on xAI and the $0.08 on ElevenLabs are not the same thing. xAI's rate covers the model's reasoning. ElevenLabs bills the LLM on top. For the full head-to-head, read our xAI voice agent vs ElevenLabs comparison. For self-built stacks, see the voice agent stack guide and best LLM for voice agents.

Where the builder helps and where it hurts

The single-model approach is strong on speed and simplicity. Fewer hand-offs mean fewer places to break. One meter makes budgeting easy.

The trade-offs are real too:

  • No LLM swap. You cannot route hard turns to a different model. Reasoning quality is xAI's.
  • Beta product. Builder features and labels can change. Pin model versions to avoid silent upgrades.
  • Opaque internals. A single model hides whether an error came from hearing, reasoning, or speaking.
  • Limited export. xAI documents no way to export a builder agent as API config. Moving later means rebuilding.

We cover the wider architecture debate in full-duplex voice agents.

Common failure modes to test

These failure modes show up on speech-to-speech agents on any platform. Each one is cheap to test and expensive to discover in production.

Failure modeWhat the caller hearsHow to catch it
Early cutoffAgent talks over a caller reading a numberScripted pauses mid-utterance
Tool overlapTwo answers play at once after a lookupSlow tool responses in test
Silent transfer failDial tone, then nothingForce a `502` on `refer`
Knowledge driftConfident wrong policy answerQuestions just outside your documents
Guardrail leakAgent reads back sensitive dataAdversarial social-engineering scripts
Language slipAgent switches language mid-callCode-switching callers
Noise misfireBackground TV triggers barge-inNoisy audio profiles

Our testing speech-to-speech voice agents guide covers each method in depth.

Testing a Grok voice agent with Evalgent

A two-minute build is a starting line. xAI's own launch post says a voice agent is easier to judge by ear than by benchmark. We agree, and we would add scale. A few calls by ear will not surface a 5% failure rate.

Evalgent tests your Grok voice agent the same way it tests any platform:

  • Scenarios reproduce noisy, accented, interrupted, and off-script calls.
  • Profiles vary caller behavior, so results split by cohort.
  • Metrics score task completion, latency, tool accuracy, and guardrail adherence against thresholds.
  • Evaluations run batches of synthetic callers against your number or SIP trunk.
  • Reviews let you hear exactly where the agent struggled.

The output is the number your budget needs: cost per resolved call on your own traffic, not a vendor benchmark. Run it before launch and after every playbook change. For the full method, see our AI voice agent testing pillar.

Frequently asked questions

What is the xAI voice agent builder?

The xAI voice agent builder, called Grok Voice Agent Builder, is a no-code tool in the xAI console. You describe calls in plain language, add documents, tools and guardrails, and pick a voice. It runs on the Grok Voice speech-to-speech model. xAI launched it in beta on July 1, 2026, and says a first agent takes about two minutes.

How much does the Grok voice agent builder cost per minute?

As of October 2026, Grok voice agents bill at $0.08 per minute of audio, with voices included and no platform fee. Telephony on a free xAI number adds $0.01 per minute. That makes the base rate about $0.09 per minute, or about $90 per 1,000 minutes. Tool calls and storage may add small amounts. Check current pricing on xAI's site.

What is Grok Voice Think Fast 2.0 pricing?

Grok Voice Think Fast 2.0, model name grok-voice-think-fast-2.0, costs $0.08 per minute of audio as of October 2026. That equals $4.80 per hour. Text inputs into a voice session are listed at $0.004 each. The same rate applies to builder agents and direct API use. Check xAI's pricing page for changes.

How do you set up a Grok voice agent to answer phone calls?

Open Voice Agents in the xAI console and pick a preset. Write the playbook, upload documents, add tools, choose a voice and set guardrails. Then attach the free xAI phone number, or route your existing number over direct SIP from Twilio, Telnyx or Plivo. Preview in the browser, test at volume, then go live.

Can you clone a voice with Grok voice?

Yes. Upload a reference clip of up to 120 seconds, ideally 90 seconds or more, and the clone works in builder agents, the realtime API and text-to-speech. Each team can create up to 30 custom voices free in the console. As of October 2026, custom voices are available only in the United States, excluding Illinois.

How does Grok voice compare to ElevenLabs?

Both list $0.08 per minute, but they meter different things. Grok's rate covers a single speech-to-speech model, including its reasoning. ElevenLabs Agents bills hosting at $0.08 per minute and charges the LLM and telephony separately. ElevenLabs offers more LLM and voice choice. Grok offers one simple meter. Test both on your own calls.

What languages and voices does Grok voice support?

xAI marketing lists 25+ languages. The speech-to-speech docs name 20+, including English, Spanish, Hindi, Arabic, Japanese and Portuguese. The model detects language automatically. xAI lists 28 built-in voices, including Ara, Eve (the default), Leo, Rex and Sal. Every voice speaks every supported language. You can also clone a custom voice.

Do you need to test a Grok voice agent before launch?

Yes. A no-code builder makes an agent easy to ship, not ready for real callers. Accents, noise, interruptions, slow tools and failed transfers break agents on every platform. Test with synthetic callers at volume and track cost per resolved call. Independent platforms such as Evalgent run these tests against your live number.

The bottom line

The Grok Voice Agent Builder costs $0.08 per audio minute plus $0.01 for a free number, which makes it one of the simplest voice agent meters available. The price that decides your budget is cost per resolved call, and only testing on your own calls reveals it.

Want that number before launch? Book a demo and let Evalgent test your Grok voice agent independently.

Related Articles