When I began building Solti's bulk audiobook publishing system, the obvious temptation was to think in clicks: upload this file, select that voice, press this button, move to the next book. That is how a person completes one title. It is not how a dependable system publishes one hundred titles, never mind one thousand.
The real system had to move finished books through Google Play Books, Google’s auto-narrated audiobook workflow, Spotify, and a wider distribution handoff. Each platform had its own account, screens, processing delays, and vocabulary. A successful click could mean submitted, processing, accepted, live, or merely “the page did not complain.” Those are not the same thing.
My most useful lesson was simple: an internal task can finish without the external result being finished. Once that distinction is taken seriously, the entire architecture changes.
The wrong starting point is the browser.
Browser automation looks like the heart of the product because it is the visible part. You watch fields fill, files upload, and buttons get pressed. It feels productive. It also encourages a dangerous shortcut: treating the absence of an error as evidence of success.
Platforms are asynchronous. A Google ebook may be submitted but not yet live. An audiobook may be processing. Spotify may accept a submission before its catalog page becomes the final source of truth. A worker can lose its connection after clicking a button and before recording the response. If it starts over blindly, it may create a duplicate.
It can perform a step and collect evidence. The platform’s current record remains authoritative for publication; the database remains authoritative for what our system intends to do next.
So I stopped treating “automate the website” as the main problem. The main problem became: preserve identity, state, evidence, and recovery across a chain of imperfect websites.
Build the ledger before you build the agents.
Our early workflow used a spreadsheet as the human-facing control surface. That was useful because every title had one visible place in the queue. But a serious publishing system needs more detail than a status cell can safely carry.
For each book, the ledger needs a stable internal identity, its source assets, exact hashes or revisions, the publishing profile, the expected account identity, platform IDs, the last action attempted, the latest observation, and the evidence time. Similar titles and filenames are not enough. At catalog scale, “probably the same book” is an invitation to publish the wrong thing into the wrong account.
We eventually separated the lanes. Google ebook, Google audiobook, Spotify, and distribution each have their own state and evidence. That prevents a submitted ebook from masquerading as an eligible audiobook, or a Spotify submission from being counted as a downstream publication.
The ledger also needs to distinguish these three questions:
- What did we ask the worker to do?
- What did the worker observe immediately afterward?
- What does the platform say now?
If those answers are stored as one status, the system will eventually lie—even if nobody intended it to.
Give agents jobs. Do not give them reality.
We created task-specific agents because the workflows genuinely differ. Arthur handles book production upstream. Sam handles Google ebooks. Sue handles Google audiobooks. Chris handles Spotify. Dan oversees publishing. The names make the operation easier to understand and let each role carry focused rules.
But names and prompts are not an operating system. Agents should not decide that a book is published because a conversation sounds optimistic. They should read the same ledger, work within an explicit profile and platform lane, and return durable evidence.
This is where a lot of “agentic” systems go wrong. They make the model the state machine. Language models are excellent at interpreting an odd rejection message or explaining what an operator should inspect. They are not the place to keep counters, enforce unique work, or remember whether a submission already happened after a restart.
Our rule became: deterministic code owns routing, eligibility, account identity, retries, and completion. Agents perform bounded tasks inside those controls.
Separate action state from platform truth.
A reliable workflow has two related but different timelines. The action timeline records queued, claimed, attempted, completed, failed, or uncertain work. The publication timeline records what the external platform currently shows: draft, submitted, processing, needs action, live, published, hidden, or rejected.
Why both? Imagine a worker presses Submit, then loses the session before saving the confirmation. The action is uncertain. The correct response is not to submit again. The correct response is to inspect the platform using the stable book and account identity, find the existing record if it exists, and reconcile the result.
That is also why “check the database every five seconds” is not platform monitoring. A fresh database read only proves that the database can read itself. Real reconciliation must visit the correct external account, inspect the correct record, save the raw observation, normalize it, and timestamp the evidence.
We learned to show both the last attempt and the last successful verification. When the worker is offline, work stays queued and the interface says so. It does not display “running” because that looks nicer.
Use AI for judgment, not bookkeeping.
AI is useful in this system, but not everywhere. It can help draft or normalize descriptions, suggest categories and keywords, interpret platform messages, and turn a technical exception into a useful operator instruction. It can also help build the system itself: generating test cases, reviewing edge conditions, and accelerating the repetitive implementation work.
Every AI-assisted publishing value still needs a stored input revision, output revision, validation result, and usage boundary. An AI-produced description is content. It is not proof that a title was accepted. An AI summary of a platform page is not a platform ID.
The boring parts should remain boring:
- unique keys prevent duplicate jobs;
- database transactions control state transitions;
- hashes prove which assets were used;
- account bindings prevent cross-profile mistakes;
- bounded retries stop transient failures becoming infinite activity;
- fresh platform observations unlock dependent work.
AI adds flexibility at the edges. Deterministic controls keep the center from moving.
Design recovery before the happy path.
The project became more reliable when we stopped asking only, “How does this succeed?” and started asking, “What happens if the computer stops here?”
A worker can fail before an upload, during an upload, after an external success, or after recording local success. Each position has a different safe recovery. A restart must reuse verified work, preserve the platform ID, and reconcile uncertain actions before replay. One bad cover or missing audio archive should block one book, not the entire catalog.
Account identity is part of recovery. We needed separate publishing profiles because two accounts can contain similar titles. A saved browser reference is not proof of an authenticated account. The worker must observe the account identity it is using and compare that with the book’s binding before making a change.
We also learned not to confuse synthetic tests with live proof. Tests can demonstrate that locking, idempotency, and status mapping work. They cannot prove that today’s platform interface, current account permissions, or external processing behaved as expected. Both kinds of evidence matter; they simply answer different questions.
A practical build order for your own system.
If I were starting again, I would build in this order:
- Define one immutable book identity. Attach every source file, platform record, and event to it.
- Model every platform as a separate lane. Write down the observed states and the evidence required to move between them.
- Create explicit publishing profiles. Bind each lane to an expected account and reject mismatches.
- Build the action ledger. Add unique job keys, leases, attempt counts, backoff, and an uncertain state before adding browser automation.
- Add one worker and one title. Complete a human-observed pilot, including restart and ambiguous-outcome recovery.
- Add platform reconciliation. Store raw observations and separate “last attempted” from “last verified.”
- Only then add specialist agents. Give each a narrow scope, structured inputs, and evidence-backed outputs.
- Scale with fixed cohorts. Prove five, then a bounded batch. Do not silently turn a successful pilot into an unlimited queue.
The system we built has passed substantial local and workflow testing, and individual live pilots have succeeded. That still does not justify claiming the entire operation is magically autonomous. Real platform execution, account sessions, processing delays, and reconciliation remain the parts that deserve the most scrutiny.
That may sound less exciting than “AI publishes your whole catalog while you sleep.” It is also much closer to how you build something you can trust with a whole catalog.
Continue the system: See how we divide the work across five audiobook publishing agents with one ledger, then use safe recovery to prevent duplicate audiobook submissions.