---
name: personal-podcast
description: “This page tells the story of The Mindful Minute, a private podcast I made for one listener: me. It is a series of short, reflective monologues, each three to eight minutes long, meant to be heard on a walk. An AI agent, Grok, and I made 567 episodes together”. Use when planning a spoken-audio or podcast project, a text-to-speech pipeline, a read-along player, or a writing guide that several agents must follow.
---

# Building a personal podcast

## Why personal podcasts?

I try to go on long walks three nights a week, and I want to keep encouraging myself to walk more. On those walks I mostly listen to video talks or songs, or I make calls. I wanted some variety: a short, reflective moment to think about something I had made, or about a topic I was curious about. That is why I made The Mindful Minute.

From the start, I defined the episodes to be reflective, and to be something I could share with my kids, so there is a safe element built in. Each episode carries a rating, which lets some episodes go further while the family-friendly ones stay safe. I am a big fan of this format: a few minutes, one voice, one idea, and then back to the walk.

## 1. The idea and the sources

### The idea

I wanted each episode to feel like a thoughtful friend thinking out loud: one voice, a small idea, a surprise along the way, and an ending that sends me back into the world, for example to look closely at the bark of a tree. The episodes would be short, private, and made from things I already care about, rather than made for everybody.

I gave the inspiration and the sources, wrote the briefs, listened to the results and made the decisions about format, rules and direction. Grok, an AI agent, wrote the episodes and built the tools under my direction, and made them for me to enjoy and to reflect on my own life: projects I started, plots I made songs about, videos I generated, and books I bought that are still waiting to be read. This is not a public feed or a commercial product.

### Where the material came from

The strongest core input was my own creative work. The episodes grew from four personal sources:

- **Videos I generated.** I generated hundreds of fun videos with an AI video generator, on a whole range of topics; the actual count is too shocking to admit. Their scenes and moods became seeds for these plots.
- **Songs I made.** I also wrote hundreds of song plots and generated songs from them with an AI music generator. From those, I hand-picked my favorites and curated them into albums, each about a specific topic. They served as inspiration for some episodes.
- **Curious conversations and projects.** Much of the material came from the conversations Grok and I have had over time: coding sessions, project ideas we discussed, and the things we built and explored together. My own background, growing up in Madurai and training as an architect and designer, gives some episodes their places and flavor, but it is never turned into a list of facts about me.
- **Books waiting to be read.** Over the years I legally bought a number of illustrated reference books, but most of them sat unread in my digital library. I wanted to use pages from them as inspiration, so the episodes became a gentle way back into those books. A small display at home now shows those pages one after another, and they became the source of the largest season.

In short, the path went from these sources into scripts, from scripts into audio, and from audio into a companion player and, later, into a plain folder of audio files for walks.

## 2. Finding the format and the rules

The first scripts were written for an AI speech generator that accepted about 5,000 characters of text at a time, which is roughly 730 to 800 words or about six minutes of speech. Many of the rules below came from testing those first episodes and listening to them.

### Monologue, not dialogue

The early batches tried both forms. One set of ten included five monologues and five two-voice dialogues, in which two speakers took turns. After listening to both, I chose **monologue only**: one voice for the whole episode. A narrator may briefly report what someone said, but there is never a second speaker. I kept the five early dialogues as they were, marked as older exceptions rather than mistakes.

### Length

The speech tool's text box became a fixed rule: every script, counting the title, the intro and the narration, must stay **under 4,900 characters**. That limit beats any word target. A few early scripts written before the rule are slightly longer and are marked to be trimmed before they are voiced again.

### Structure and spoken sign-off

Each episode has a title, a short **Intro**, and the main **Narrative**. The intro opens with one plain sentence that names its maker: "Made by", followed by the name of the AI agent that wrote the episode. The sign-off is said once and never repeated in the narrative. At first the rule fixed one agent's name; later I asked for it to be generalized so that any agent that writes an episode signs with its own name, which keeps authorship honest if several agents contribute.

### Rotating audiences

I asked for episodes written for different imagined listeners in turn: Indian families living abroad, someone from Madurai, someone who knows Bangalore, a Californian who knows India, and people with no link to India at all. The audience is never named aloud; phrases such as "for those of you who..." were removed. It is recorded only as metadata. About a third of the episodes are suitable for a family walk with children of about ten and fourteen, and Madurai is kept as a main setting in only a small share of episodes.

### No names

Spoken text never names a company, brand, book, author, publisher, franchise or fictional character. Historical figures appear only where they are essential to a fact. My children and home address are never mentioned, living private people are never named, and nobody is given an invented quotation. Facts must be true, and fiction must be clearly fiction.

### Content ratings

Each episode carries one of four ratings: **All** (for everyone), **T** (teen), **E** (strong language) and **Adult only**. E and Adult-only episodes use real slang and swearing the way people actually talk, and their intros carry a short spoken warning that is worded freshly each time. Episodes rated All or T carry no warning.

### The repeated-line lesson

The most useful rule came from a mistake. When a script came out short, the writing process filled the gap with a stock sentence. In three episodes, the line "Be gentle with your pace. These episodes are companions, not tests." appeared between 24 and 29 times each, and in two others "Momentum loves small doors." appeared six to eight times. I heard it immediately and flagged it. Grok cleaned up the episodes, a checker now looks for repeated sentences within an episode and across episodes, and the rule is simple: **if a script is short, end it; never pad**. A few smaller repeats across episodes, such as "So here is the walk game.", are still listed in the project's records as defects to fix.

## 3. The seasons

![The four seasons, drawn to scale by episode count.](shots/03-seasons.svg)

### Season one: 59 episodes

The first episodes were written on the first two days and tested many things at once: monologues and dialogues, topics from technology to music, and two versions of each script, one with bracketed sound cues for a music-and-speech tool and one plain. They were written before most of the rules existed, so they carry no spoken sign-off and a few of them break later rules. I kept them as they are, marked as older exceptions.

### Season two: 91 episodes

For the second season I gave Grok a free choice of topics, with the songs I made and the videos I generated as inspiration. When Grok reviewed its own work honestly, it found the season too inward-looking: too many episodes leaned on my own songs and videos, the same devotional and slang threads came back again and again, and too many stories pulled toward Madurai. Instead of deleting the season, I kept it as a backup and asked for a different direction for the next one.

### Season three: 100 episodes

The third season answered that critique with short fiction, which we called "small nuggets": stories about three-dimensional shapes, lecture-style journeys, the pleasure of being briefly scared, and small human moments. Every plot carries its own twist, so the story itself is the surprise. Episodes next to each other never share a form, and the forms vary widely: mock lecture, field guide, fable, letter, mystery and confession. Examples include *You Have Been Tying Your Shoes Wrong*, a mock lecture that ends with the speaker's own shoe coming untied, and *The Pillars That Sing*, a quiet mystery about stone pillars that ring when tapped.

### Season four: 317 episodes

The fourth and largest season, which I approved on the third day, was written for a display at home that shows pages from inspiring books, one spread at a time. Each episode is inspired by one of those pages and is meant to nudge me to pick up a book of that kind and read it. The rules for this season were strict:

- **Abstract sourcing.** The episode never names the book, author, publisher, series or any character, and it never quotes or closely paraphrases the printed page. The page text is used only to check facts.
- **A quiet hook.** Each episode suggests a kind of book, a way of looking or a place, but never a title and never a hard sell.
- **One line for the display.** A single sentence points at the page on the screen, and the episode still makes sense on a walk with no picture.
- **Themes, not plots.** I decided to include pages from comics and graphic novels; they are used for their themes, such as robots or dragons across cultures, without retelling a plot or copying dialogue.
- **Ratings.** 266 episodes are rated All, 38 are T, 11 are E and 2 are Adult only.

An example is *The Tallest Living Things*, which looks at a trunk rising out of the frame into mist and ends with the listener finding the top of that trunk on the screen.

## 4. The companion player and the voice

### What the player had to be

Before Grok built the player, we studied how listening apps lay out reading and playback, and I thought about what a calm companion app should feel like on a walk: large, readable text, one episode in focus, quiet controls, and nothing that asks for attention. The result is a single web page that works offline on a phone. It has:

- **Read-along highlighting:** the words light up as they are spoken, and tapping a word jumps the audio to it.
- **An episode browser** grouped in folders of 50, with search by title or episode number, and a resume point for each episode.
- **A contents sheet** that shows each episode's intro and narrative as chapters.
- **A bottom player bar** with skip, speed and next-episode controls.
- **A passcode screen.** The page is private, so I chose a simple passcode screen in front of it, and it asks search engines not to index it. It is a light gate, not real security.

The first version, **v0.01**, went online on Oct 8 behind that passcode screen, with the texts of the first 250 episodes and audio for five of them. The local version on my home computer now lists all 567 episodes, and 104 of them have audio. Filters by audience, place, rating and season have been drafted, with sixteen tags, but they are not in the player yet.

<!-- gallery: The companion player at phone width (local build) -->
![An episode open in the reader: title, running time and the start of the Intro.](shots/04a-player-episode.jpg)
![Read-along: the words already spoken are dark and the rest are pale, so I can follow along.](shots/04b-player-read-along.jpg)
![The contents sheet shows each episode's Intro and Narrative as chapters, with a link to browse all episodes.](shots/04c-player-contents.jpg)
![The episode browser: folders of 50, search, running times, and a resume point for a half-heard episode.](shots/04d-player-browser.jpg)
<!-- /gallery -->

### Choosing a voice

The scripts were first written for an AI speech generator, but its text box and its handling of sound cues were uncertain, and I wanted to own the voice. Grok ran a small pilot: three open-source text-to-speech models each read the same episode and a one-minute sample, which gave six recordings compared on a simple scorecard. One model, a small open-source model with about 82 million parameters and an American English male voice, was chosen, and a second was kept as a backup. I plan to add a commercial voice later as a second version of each episode, and the file structure leaves room for it (see section 5).

## 5. Tech learnings

![The production path for one episode.](shots/05-pipeline.svg)

### Read-along timing

To highlight words as they are spoken, the player needs to know when each word starts and ends. An open-source speech-recognition model listens to the finished audio and writes those times into a small data file. The recognized words are then lined up with the actual script, so the text on screen is always the written text; between 94 and 100 percent of the words matched. In a later test, the voice model's own timing output matched the script exactly, which could remove the recognition step and save about 45 seconds per episode.

### Cue markers, and why separate files replaced them

The first tools returned one long audio file per episode, so we needed to find where the intro ended. The solution was a spoken marker: the word "Ding." at the end of the intro and again at the end of the narration, found later by the recognition model. It worked, but the recognition model missed 18 of 49 short intro markers, and a recovery step had to find them. Once the voice ran on my own computer, I questioned whether the markers were needed at all and suggested the better answer: **voice the intro and the narration as two separate files** and join them with a fixed pause of about two seconds. The boundary is then exact, no marker is needed, and nothing has to be searched for. The markers survive only in versions meant for tools that return a single file.

### Moving production to my home computer

The first 50 episodes were voiced on a shared cloud computer that Grok works on. It took about 170 seconds to voice and 92 seconds to time each episode, only two voices could run at once because of memory, one run was killed for lack of memory, and the machine froze for about 36 minutes; a batch of 49 episodes took about two hours. Moving the work to my home desktop computer with a graphics processor changed the numbers: about 26 seconds to voice and 45 seconds to time a five-minute episode, about 70 to 80 seconds in total, or roughly one hour for a batch of 50. The remaining episodes need about nine to ten hours and about 2.7 GB of disk space, so I cleared disk space first, from 18 GB free to 28 GB.

One practical lesson: a long job started from a remote session stopped when that session closed. Long jobs now run under the computer's own job scheduler, which keeps them going after the terminal closes and keeps the computer awake.

### Batch jobs and a progress view

The remaining work is split into twelve batches, eleven of 50 episodes and one of 17. They run one at a time, so two voice jobs never compete for the graphics processor. Finished episodes are skipped, so an interrupted batch can simply be started again. I start the batches myself with one command. Progress shows as bars in the terminal and on a small status page that refreshes every five seconds, so I don't need to read log files, and the player page rebuilds itself after each batch.

![The status page: one row per batch, with finished, partial and waiting batches.](shots/06-status-page.jpg)

### Per-episode folders

At first the scripts lived in large documents of 50 episodes each, which was convenient for my listening tests but awkward for production. I asked for one folder per episode, with number-first file names, Markdown sources and a walk folder. Now each episode has its own folder, and the documents of 50 can be generated from the folders whenever they are needed.

![One folder per episode. I edit the Markdown files; everything else is generated.](shots/07-episode-folder.svg)

- **Markdown for anything a person edits:** the script, the version with sound cues, a music and style prompt written without naming any tool, and notes on tone, pace and mood.
- **Plain text for anything spoken.** The voice reads every character literally, so a heading symbol could be spoken aloud. The intro and narration text files are generated from the Markdown script and are never edited by hand.
- **Number-first names.** Every file starts with the three-digit episode number, so files sort correctly and a stray file still says which episode it belongs to. Numbers are permanent: never reused or renumbered.
- **Room for a second voice.** Each voice writes its own audio and timing pair, so a later commercial voice adds files without overwriting anything.
- **A walk folder (planned).** Finished audio with readable names and track information, for any music app on a phone, with no web page needed.
- **Protection for hand edits.** Splitting the scripts again only overwrites files that nobody has changed.

### One source of truth and an activity log

Because several agents may work on the project, and any one of them may run out of credits, I asked for a single source-of-truth document and a separate activity log. The document records the current state, the yes-or-no rules, the softer guidance on voice and feel with real examples, notes on each season, a dated history of decisions and reversals, and a copy-paste runbook. Older rules are marked as replaced rather than deleted, so nobody revives them by accident. The activity log records every working session: which agent worked, what I asked, what was done, which files changed, which checks ran and what is still open. Past entries are never rewritten.

### Lessons learned

- **Listen early.** The repeated-line bug and the dialogue question were both settled by listening, not by reading.
- **Never pad.** A short episode is better than a filled one.
- **Write rules as yes-or-no checks** where possible, and keep the softer guidance separate, with examples.
- **Critique a whole batch honestly** before writing the next one; the second season's self-critique shaped everything after it.
- **Design files for the machine that reads them:** Markdown for people, plain text for the voice.
- **Separate files beat markers.** When you control production, make the structure exact instead of detecting it afterwards.
- **Keep records that outlive the conversation.** A single source of truth and an activity log let another agent continue the work.

### What's next

- Voice the remaining batches.
- Build the walk folder of tagged audio files.
- Add the audience, rating and season filters to the player.
- Decide whether the voice model's own timing replaces the recognition step.
- Fix audio seeking on the hosted version, which does not yet support jumping within an audio file on every phone.
- Publish an updated version of the player, once I approve it.

## Related projects

- **[Mindful Minute](https://mindful-minute.pages.dev/)**: the companion player itself, a private read-along page behind a passcode screen.
- **[AI Video Gallery](https://briji-projects.pages.dev/)**: the gallery of the videos I generated, one of the core sources for the episodes.
- **[Thara Local Studio](https://thara-local-studio.pages.dev/)**: the albums of songs I made, another core source.

## Credits

Written by me with Grok, an AI agent. I gave the inspiration, the sources and the decisions; Grok drafted the episodes, built the tools and drafted this page under my direction. October 2026.
