05Writing

TalkToMe: I made my coding agent call me

· voice · agents

Why open another app when the agent already working in your project can just call you?

I use coding agents for a lot of my work now, and one thing that has always felt slightly weird is how much context they have compared to how limited the interface still is. My agent can have my repository open, know what we changed ten minutes ago, remember a long conversation, run commands, inspect files and generally be pretty deeply involved in whatever I’m building.

But if I want to talk through an idea with it, I still have to go find that exact session and start typing.

So I built TalkToMe. You tell the agent to call you, your Mac rings, you answer, and you can talk to the same agent session that was already working on the project.

That last part is the bit I actually care about.

The agent already has the context

There are already quite a few ways to talk to an LLM. I didn’t really want to build another one.

If I have been working with Codex for an hour and it has already explored the codebase, edited files and discussed an implementation with me, it feels silly to open a separate voice assistant and explain the entire thing again. The coding agent already has everything it needs; I just want another way to communicate with it.

TalkToMe therefore doesn’t provide an LLM at all. The agent keeps its existing model, tools, files and conversation history. TalkToMe handles the microphone, transcription, call UI and speaking the response back to you.

The flow is basically:

  1. I tell the agent, “Call me.” or i mention a condition like “call if you are not sure which way to proceed.”
  2. It runs the TalkToMe command with its session ID.
  3. A call appears on my Mac.
  4. I answer and start talking.
  5. What I say gets sent back into the existing agent session, and its response is spoken back to me.

Once the call is over, the session is still exactly where it was. There isn’t a new conversation to merge back into your work or another chat history to keep track of.

The app itself mostly sits in the menu bar when you aren’t using it.

Why voice?

I definitely don’t think I want to code entirely through voice. If I know precisely what I want the agent to do, typing something like “move this into the service, update the tests and don’t touch the public API” is probably quicker and clearer.

Where voice becomes useful is when I don’t know exactly what I want yet.

A lot of programming is me thinking something like, “Okay, maybe we put this in the worker… but then the websocket state becomes annoying… unless we keep that part in the main process… actually wait, could we just…”

That is a surprisingly annoying thing to type.

Talking is much better for that sort of messy thinking. I can reason through an architecture, explain why I dislike an implementation, or throw half-formed ideas at the agent without first converting all of that into a clean prompt.

In other words, I’m not really using voice to tell the agent what to code. I’m often using it to figure out what I want to tell the agent to code.

Making it feel like a conversation

Another thing I didn’t want was push-to-talk. Having to hold a button, speak a perfectly formed thought and then release it starts feeling more like voice commands than an actual conversation.

TalkToMe uses Smart Turn, a small local turn-detection model, to decide when you’ve probably finished speaking. This means a pause in the middle of a thought doesn’t necessarily immediately send the message.

There is also a fixed-pause mode if you want something simpler and more predictable.

This sounds like a relatively minor part of the app, but it makes quite a big difference in practice. When you are thinking out loud, you pause a lot more than you realize.

Local speech

I also wanted it to be possible to run most of the voice pipeline locally.

Whisper can handle speech-to-text, while the agent’s reply can use the macOS system voice or Kokoro for local text-to-speech. Once the models are downloaded, they can run offline.

If you want better cloud voices, you can use ElevenLabs instead for speech recognition or TTS. In that case, obviously, the audio or response text needed for whichever service you enable gets sent to ElevenLabs.

The actual intelligence still comes from your coding agent either way. TalkToMe is really just sitting between you and an existing agent session.

Agent support

Codex is the integration I’ve been using most directly, but TalkToMe also has connection paths for Claude Code, Hermes Agent, OpenClaw and generic agents that can run shell commands.

This was actually one of the more annoying parts to design because different agents expose their sessions differently. Codex, for example, can be handled differently from an agent where TalkToMe has to cooperate through commands to listen for messages and return replies.

I ended up keeping that complexity behind the connection layer so the experience on my side remains the same: I ask whichever agent I’m already using to call me.

There is also an experimental remote bridge for cases where the agent isn’t actually running on your Mac. The idea is that an agent working on a server can still ring your laptop while the microphone audio itself remains local.

That part is still experimental though.

Where I want to take this

This is an early Mac app, and right now the experience I care most about is pretty specific: I’m working with an agent on my Mac, it already understands the project, and I want to talk through something without moving to another interface or losing that context.

But building it has made me think a bit differently about agent interfaces in general.

Almost every agent today lives somewhere you have to go to. You open your terminal, IDE, browser or chat window and interact with it there. That makes sense when the agent is basically waiting for instructions.

It gets more interesting as agents start doing longer-running work.

If an agent is already working independently and eventually gets to something where it needs a decision, maybe I shouldn’t have to keep checking its terminal. It could message me. It could send a notification. Or, for something complicated enough that a back-and-forth would be useful, it could just call.

I don’t think voice replaces the terminal or chat. I think it becomes another interface an agent can choose when it makes sense.

For now though, I mostly just like being able to type:

Call me.

…and have my Mac actually ring.

TalkToMe is open source here:
github.com/rohanprichard/talktome

It’s still early, so if you use Codex, Claude Code or another coding agent and try it out, I’d especially like to know where the call experience feels useful, annoying, slow or god forbid, actually like it.