I Put an AI on My Phone Line. Then Another AI Called It.

Posted on Wed 30 September 2026 in Technology

I Put an AI on My Phone Line. Then Another AI Called It.

This project started because I wanted to play with voice again.

I have not done much with voice systems in years, and I thought it would be fun to combine some old-school telephony with newer local AI tools. What started as a simple experiment turned into a complete AI-powered IVR that can answer a real telephone call, listen to the caller, transcribe what they say, decide how to respond, generate speech, and send that audio back through the phone system.

Then something happened that I absolutely did not plan for.

Another AI called it.

And, because apparently I cannot leave well enough alone, my AI happened to sound like Duke Nukem.

Yes. That Duke Nukem.

The bubble-gum-chewing, alien-kicking video game hero of the 1990s.

Before I get to that part, though, it probably helps to explain what I actually built.

IVR Voice AI Call Flow

The Call Flow

The actual path of a call looks like this:

Caller
  |
  v
Xfinity Cloud / PSTN
  |
  v
Cisco ISR4351 Router
  |
  | FXO
  v
SIP Trunk
  |
  v
Asterisk 20.6.0
  |
  v
IVR Orchestrator
  |
  v
Faster-Whisper
  |
  v
Qwen 3.5 9B via Ollama
  |
  v
Coqui XTTS-v2
  |
  v
IVR Orchestrator
  |
  v
Asterisk
  |
  v
Caller

There is something I really like about this architecture.

At one end, I have the public telephone network and an analog FXO interface. That is technology with roots going back decades.

At the other end, I have local speech recognition, a large language model, and neural text-to-speech.

A telephone call comes in through Xfinity, crosses the PSTN, hits a Cisco router over FXO, gets converted into SIP, lands in Asterisk, gets handed to an application, becomes text, gets processed by an LLM, becomes speech again, and then travels all the way back out to the caller.

That is a ridiculous amount of technology involved just to say, "Hello."

And I love it.

Not Everything in the System Is AI

One thing I wanted to keep very clear while building this is that the entire system is not "AI."

There are really three AI workloads in the call path:

  1. Faster-Whisper handles speech-to-text.
  2. Qwen 3.5 9B handles the conversational response.
  3. Coqui XTTS-v2 turns the generated response back into speech.

Everything else is infrastructure or application logic.

Asterisk is not AI.

The Cisco router is not AI.

The PSTN is definitely not AI.

Ollama is not the language model itself. It is the runtime serving the model.

And the IVR orchestrator is intentionally not AI either.

That part matters.

The Orchestrator Is the Traffic Cop

The IVR orchestrator is the piece that ties the whole thing together.

It coordinates the live call between Asterisk, speech recognition, the language model, and text-to-speech.

Its job is to manage things such as:

  • receiving audio from Asterisk
  • deciding when the caller has finished speaking
  • sending audio to Faster-Whisper
  • receiving the transcription
  • maintaining conversation state
  • sending the appropriate context to Qwen
  • enforcing rules that should not be left to the language model
  • sending the response to XTTS
  • receiving generated speech
  • handing that audio back to Asterisk
  • managing errors and timeouts

That last part is important.

There are some things I do not want to depend on the language model remembering.

For example, the system is programmed to identify the caller and provide a recording disclosure.

Those are deterministic application rules.

The LLM can handle the conversation, but the application decides what absolutely has to happen.

That separation turned out to be one of the more important lessons from the project.

Why I Needed More Than One Computer

Running a model is one thing.

Running several models quickly enough to hold a telephone conversation is something else.

Real-time voice has a deadline.

If someone says something and your system takes fifteen seconds to answer, the technology may technically be working, but the conversation is not.

Humans notice delay very quickly.

The system currently spreads the workload across several machines.

| Host | Role | |---|---| | voicertr | Cisco ISR4351 handling the Xfinity/PSTN side | | astricks | Asterisk 20.6.0 | | ai | IVR orchestrator, Faster-Whisper, Ollama, and Qwen | | babyAI | Coqui XTTS-v2 speech generation using CUDA |

This was not an attempt to build a cluster because clusters are cool.

It happened because the workload demanded it.

Speech recognition, LLM inference, and speech generation all have different compute characteristics. Once they start competing for the same CPU, GPU, memory, and storage resources, latency gets ugly very quickly.

Splitting the workloads let me keep the conversation moving.

Speech-to-Text: Faster-Whisper

The first AI component sees the caller's audio.

Faster-Whisper transcribes that audio into text.

My current configuration uses the tiny.en model with CPU inference and INT8 compute.

That might sound like the least exciting part of the system, but speech recognition is where the whole thing can fall apart immediately.

Telephone audio is not clean studio audio.

There can be:

  • background noise
  • compression
  • clipping
  • accents
  • people talking too quickly
  • people mumbling
  • people interrupting
  • silence that may or may not actually mean the caller is done

If the transcription is wrong, everything after it is working with bad input.

Garbage in, extremely confident AI garbage out.

The Language Model: Qwen Through Ollama

Once the caller's speech becomes text, the orchestrator sends the appropriate context to Qwen 3.5 9B through Ollama.

The model generates the response text.

This is where the conversational behavior lives.

But I discovered pretty quickly that a telephone conversation is very different from a normal chatbot session.

A chatbot can generate three paragraphs and nobody cares.

A voice IVR that answers every question with three paragraphs becomes unbearable almost immediately.

The responses need to be short.

They need to sound conversational.

They need to avoid rambling.

And the system needs enough context to understand the conversation without feeding it an ever-growing wall of text on every turn.

That is still one of the areas I continue to tune.

Text-to-Speech: Coqui XTTS-v2

Once Qwen generates a response, the text goes to Coqui XTTS-v2.

XTTS turns the response into speech.

This is running on a separate machine, babyAI, using CUDA.

And this is where I made a completely unnecessary but highly entertaining decision.

I cloned a voice modeled after Duke Nukem.

If you were around PC gaming in the 1990s, you probably know exactly the voice I mean.

For anyone who does not: imagine every exaggerated action-movie hero rolled into one person, given a microphone, and then told subtlety was illegal.

That was now my IVR.

And Then Rocket Mortgage Started Calling

Here is where the project became much more entertaining than I expected.

I started getting repeated calls that appeared to be from Rocket Mortgage.

I had not contacted them in years.

My mortgage is not even through them.

But the calls kept coming.

Eventually one of those calls landed on the AI IVR.

And the caller appeared to be an automated voice agent as well.

So, without intending to, I had created:

AI versus AI over the public telephone network.

My IVR was programmed to make sure it obtained the caller's name and provided the recording disclosure.

That sounds simple.

It is the sort of thing a human caller would barely notice.

But the automated caller seemed to have its own expected conversational flow.

At one point, my IVR asked for the caller's name in the middle of that flow.

And that was apparently enough to confuse the other system.

The conversation started going sideways.

My AI wanted a name.

The other system wanted to continue whatever script it was following.

Neither side behaved exactly the way the other expected.

And suddenly I was listening to two automated systems trying to negotiate control of a telephone conversation.

One of them sounded like Duke Nukem.

It was glorious.

When Two State Machines Walk Into a Phone Call

The funny part is that underneath the AI, there is a legitimate engineering lesson here.

Conversational systems still have state.

One system may expect something like:

INTRODUCTION
    ->
SALES MESSAGE
    ->
QUESTION
    ->
EXPECTED RESPONSE
    ->
NEXT QUESTION

My system has its own state:

ANSWER CALL
    ->
PROVIDE DISCLOSURE
    ->
IDENTIFY CALLER
    ->
LISTEN
    ->
TRANSCRIBE
    ->
GENERATE RESPONSE
    ->
SPEAK

Now put those two systems on the same telephone call.

Each one has goals.

Each one has assumptions about turn-taking.

Each one may be waiting for a specific kind of response.

Each one may have timeout rules.

Each one may attempt to recover when the conversation does something unexpected.

A human can usually handle an interruption such as:

"Before we continue, what is your name?"

Another automated agent may not.

That is exactly what made the call so interesting.

It was not simply two language models having a conversation.

It was two complete conversational systems, each with its own logic and expectations, colliding over a normal telephone call.

The Hard Part Is Latency

One of the biggest lessons from the project has been that conversational AI is not just about model quality.

It is about time.

The complete path looks roughly like this:

Caller speaks
    |
    v
Detect end of speech
    |
    v
Transcribe audio
    |
    v
Generate LLM response
    |
    v
Generate speech
    |
    v
Send audio back through Asterisk
    |
    v
Caller hears response

Every stage adds latency.

Even if every individual component is "fast," the delays add together.

There is also a difference between:

The model can run on this computer.

and:

The model can run fast enough on this computer to participate in a telephone conversation.

Those are not the same requirement.

This is one of the main reasons I ended up distributing the workload.

Things That Do Not Work Perfectly

This project has also been a reminder that demos are easy and systems are hard.

There are plenty of edge cases.

A few examples:

  • the caller talks over the AI
  • speech detection decides the caller is finished when they are not
  • Whisper mishears a word
  • the LLM generates too much text
  • TTS takes too long
  • audio levels are wrong
  • background noise creates nonsense transcription
  • one model consumes resources another model needs
  • the caller says something completely outside the expected flow
  • the LLM gives a response that sounds fine in text but terrible when spoken aloud
  • silence lasts just long enough to make the system wonder whether the call is over

None of these problems are individually shocking.

The interesting part is that they all happen in the same pipeline.

The Old Stuff Still Matters

One of the things I enjoyed most about this project was getting back into voice.

It also reminded me that modern AI does not replace traditional infrastructure knowledge.

You still need to understand:

  • FXO
  • SIP
  • RTP
  • codecs
  • Asterisk
  • routing
  • Linux
  • application state
  • audio handling
  • networking
  • CPU and GPU resource constraints

The AI models are only one part of the system.

If the SIP trunk is broken, Qwen cannot save you.

If the FXO interface is not working, Whisper is irrelevant.

If Asterisk cannot move the audio correctly, your fancy cloned voice is just a very expensive WAV file sitting on disk.

That is probably what I like most about the project.

It crosses a lot of different areas.

Old-school telephony.

Networking.

Linux.

Python.

Local AI.

GPUs.

Speech recognition.

Voice cloning.

And just enough ridiculousness to make the entire thing worth building.

So What Did I Actually Build?

In the end, this is not just a chatbot connected to a telephone.

It is a complete real-time voice pipeline:

PSTN
  ->
FXO
  ->
SIP
  ->
Asterisk
  ->
Application Logic
  ->
Speech Recognition
  ->
Language Model
  ->
Speech Synthesis
  ->
Asterisk
  ->
PSTN

And every part of that chain has to work quickly enough that the person on the other end still believes they are participating in a conversation.

The fact that another AI eventually called it was just a bonus.

The fact that the other AI got confused by mine asking its name was even better.

And the fact that the whole conversation happened in a cloned Duke Nukem-style voice?

That was just good engineering judgment.

Probably.


This post describes my own lab setup and my observations of calls received by the system. References to third-party companies or automated callers describe what I observed during those calls and are not claims about the internal design of those companies' systems.