As I was leaving for a hike this weekend, my wife asked where I was going.
“The Grand Coulee Dam.”
She observed that this was a four-hour drive.
I told her I had eight chapters to revise.
You see, it takes about an hour to revise a chapter while driving. Four hours out, four hours back. That’s the whole manuscript.
Press enter or click to view image in full size
Editing has always been my least favorite part of writing. You get excited about a topic, you dump everything into a draft, and then the bill comes due: eight or ten rereads before you’re comfortable. The writing is the fun part. The editing is the “tax,” and I’ve spent years of my life reading and rereading and rereading while working on books.
I hate editing and reviewing text. Print it out. Change the font. Read it aloud. They all help a little, and they all fail the same way: you’re sitting still, staring at a screen, and your brain fills in what you meant to write. My ADHD makes that worse. After the third pass, I’m reading the news.
A few months ago, I stumbled into something better. Driving to a meeting with a draft I needed to review, I sent it to an agent that reads text back with text-to-speech. When something sounded wrong, I recorded a voice note. “That last sentence was hateful, embarrassing AI nonsense. Change it to…” The agent transcribes the note, applies the edit, and reads the updated section back. Forty minutes. Two full passes on a 3,000-word document. Better draft than I left with.
When I write books, I use that same workflow to rewrite almost every word (sometimes twice), and it’s a much faster process than typing.
Now that loop is how I work — including the drive to Grand Coulee. When I create, review, and edit content, it’s almost entirely transcription and TTS. I still type some final polish at the end of a cycle. Everything before that — drafting, listening, revising — is voice.
Want to make one thing clear: I try not to edit while I’m hiking, and, for the record, I was driving to Umatilla Rock at Dry Falls, not the dam itself. And, it was also an awful day for a hike because it was 92 F.
Talking is faster than typing
There’s something about talking that lends itself to faster expression. When I’m creating, I talk. The sentence doesn’t have to survive my fingers before it exists. I can argue with myself out loud, try three phrasings in thirty seconds, and keep the one that lands. Dictation isn’t a gimmick for hands-free convenience. It’s a higher-bandwidth way to get ideas out before they cool off.
When I drive, I edit. When I’m at work, I find a conference room and talk to my computer. If I need to revise, I listen. If I need to rethink an argument, I argue it out loud and drive the revisions through audio.
I’m not using consumer “voice mode” in Claude or ChatGPT. Those interfaces usually put you on a smaller, latency-optimized model. I route TTS and transcription through an agent pointed at a real model with write access to the draft. There’s lag. That’s fine. Direction and execution beat a chatty voice toy.
I tend to prefer the Onyx voice from OpenAI's tts-1, paired with an affordable model. I’m providing most of the direction when I’m editing. In the rare case that I’m asking for assistance or complicated draft revisions, I might jump up to a more expensive model, but my edits are usually me getting rid of AI slop, and you don’t need reasoning models for that.
And there’s a lot of AI slop. AI sucks at writing.
Emotional direction, verbatim cuts
Listening changes what I notice. Reading, I skim past flat sentences. Hearing them, I hear the dead air. On a voice pass, I can give emotional direction the way a director talks to an actor: tighter here, angrier there; this paragraph is hedging — stop hedging. That kind of steering is awkward to type and natural to say.
And a lot of the time I’m not steering at all. I’m dictating the fix. Verbatim edits. “Replace that sentence with: …” Exact wording. That’s one of the main ways I unsloppify the crap AI generates — I rarely ask it to “make it better.” I hear the slop, say the real sentence, and make the agent put my words in, and what’s useful is that the process will address grammar issues automatically.
The loop is simple: TTS reads the draft, I send a voice note, the agent transcribes, edits, and reads back. Not a conversation. Direction and execution, and I’m able to work much faster with this interaction pattern.
Return to the Office or the Driver’s Seat?
Some days the desk wins. Some days the car does. I can’t always predict which. On the voice days, I get ten times more done walking or driving than I would staring at a screen. The overhead for TTS and transcription is trivial — a few dollars a day.
All of this is incompatible with the current “return-to-office” theater.
Yes, it’s good to sit next to people and experience small talk about how bad the smoke is lately. But mandating open-plan desks for people who prefer a highway and a hike optimizes for visibility, not output. If you’re forcing someone to drive in so they can sit down and do work they could have finished better elsewhere, you haven’t solved a productivity problem. You’ve created a new one that isn’t adapting to emerging interaction patterns.
Ok, sure, I’m not saying I’d drive to the Grand Coulee every day instead of driving to the office, but there are days when I rush to the office just to reserve a conference room and talk to the laptop for several hours.
The office wasn’t built for this new interaction pattern.