When I began turning Solti's managed bulk audiobook publishing service into a system, one large publishing agent seemed attractive. Give it a book, give it browser access, and tell it to keep going until the title is everywhere. It is a wonderfully short product description. It is also a poor operating model.

Publishing a finished catalog is not one job. It is a chain of account-bound jobs: preparing the Google ebook record, creating or validating the Google audiobook, submitting the same title to Spotify, handing it to wider distribution, and checking whether each platform actually made it live. Every lane has different inputs, controls, delays, and failure states.

We split that work into specialist roles. The split made the system easier to reason about, but it created a new risk: if every agent kept its own memory of a book, we had not built a publishing operation. We had built a committee.

One giant agent was the obvious bad idea.

A single agent has one apparent advantage: there are no formal handoffs. It can carry context from the beginning to the end. The trouble is that the context becomes both enormous and unreliable. Instructions for one platform can leak into another. A transient Spotify problem can obscure a clean Google result. Recovery after a restart depends on reconstructing what the agent thinks it did.

There is also no sensible blast radius. If one long-running session goes wrong, every platform lane for that book may be affected. At catalog scale, that is how a small ambiguity turns into an afternoon of archaeology.

Specialize the work, centralize the facts.

Agents may own different tasks, but they must read and write the same durable book identity, account binding, action history, platform IDs, and evidence.

Specialization gives us clearer permissions, smaller prompts, narrower tests, and more useful failure messages. The cost is that handoffs have to be designed instead of merely hoped for. That cost is worth paying.

Define agents by platform responsibility.

Our useful division is four execution roles plus oversight. The Google Ebook Publisher prepares and verifies the ebook record required by the next lane. The Google Audiobook Publisher works from an eligible live ebook and handles the auto-narrated audiobook workflow. The Spotify Publisher submits the validated audiobook to Spotify. The INaudio Publisher handles the wider distribution lane. Publishing Oversight watches the collection of lanes and calls attention to stale, blocked, or contradictory evidence.

The names are deliberately about work, not personalities. An operator should know what an agent is allowed to change simply by seeing its role. Each role receives a book ID, a publishing profile, an account binding, a lane, and a bounded instruction. It does not receive permission to tidy up unrelated titles because it happened to notice them.

That division also makes platform differences honest. Google ebook publication is not Spotify audiobook publication. A successful Google step can make a later lane eligible, but it cannot mark that later lane complete. Every agent owns its action; no agent owns reality.

One ledger is the source of operational truth.

The shared ledger is not a chat transcript. It is structured state. For every title it records an immutable internal book ID, the source asset revision, its publishing profile and expected account, each platform lane, any stable external ID, the latest intended action, the latest observed result, and when that evidence was collected.

We keep action state separate from publication state. An action can be queued, claimed, attempted, completed, failed, or uncertain. The platform record can be draft, submitted, processing, needs action, live, or rejected. Those timelines influence one another, but combining them into a single “status” field erases exactly the information needed when something fails.

This is the central lesson from the first field note about why “done” is not a publishing status. The database is authoritative about what our operation intends and what it recorded. Google, Spotify, and the distributor remain authoritative about whether a title is actually published there.

A good ledger lets a new worker resume without inheriting the old worker’s conversation. It can see that a Spotify submission was attempted, that the outcome is uncertain, and that the next permitted action is reconciliation—not another submission.

Conversation is not execution.

We wanted the agents to be visible. An operator should be able to ask what is happening, understand a rejection, or request a bounded run. That makes a conversational interface useful, but it creates a subtle temptation: allowing the conversation to become the control plane.

I do not want “I’m on it” to mean a job exists. I want a durable job row with a unique key, an allowed transition, a worker lease, an attempt number, and a book-and-profile scope. The agent can explain that record in plain English. It cannot replace it with optimism.

The same boundary applies to AI. A model is good at interpreting a strange platform message, suggesting how to repair metadata, or summarizing why a title is blocked. Deterministic code should still own eligibility, identity, routing, uniqueness, backoff, and completion. Language is flexible; state transitions should not be.

Make every handoff evidence-based.

A handoff is not “the previous agent says it finished.” It is a gate supported by a fresh observation. The Google audiobook lane should not begin because the ebook worker completed its browser script. It should begin when the correct Google account shows the correct ebook in the required state, with the observation saved against the same internal book.

The handoff packet needs to be small and explicit: book ID, profile and account, upstream platform ID, relevant asset revision, the observed upstream state, the evidence time, and the exact downstream action now permitted. If one item is missing or contradictory, the next role waits.

This feels slower than letting agents improvise. It is faster than investigating the wrong title in the wrong account after fifty books have moved through the queue.

Design the Team view for the operator.

A Team screen should reveal the machinery, not decorate it. For each role we show whether its worker is online, stopped, or unavailable; which profile and lane it controls; the current book; the latest durable activity; and the reason work is queued or blocked.

If the executor is offline, the interface must say so. If a job is merely queued, it must not look active. If an adapter is not connected, a friendly agent card cannot imply that it is publishing. These distinctions can make a dashboard feel less exciting. They also make it useful.

The operator needs profile-level start and stop controls because accounts and batches are operational boundaries. Pausing one profile should not silently pause another, and stopping an agent should not rewrite the platform state of books it previously handled.

A practical blueprint for publishing agents.

If you are building your own multi-agent publishing system, I would use this sequence:

  1. Map the real lanes. Separate jobs where permissions, account identity, platform truth, or recovery behavior differ.
  2. Create one stable book identity. Every agent, asset, event, and platform record must point back to it.
  3. Define each role’s contract. Specify its inputs, permitted actions, evidence output, and stop conditions.
  4. Build the ledger and transitions. Do this before giving any agent a browser.
  5. Bind work to a profile and account. Verify the observed account before any external mutation.
  6. Require evidence at handoffs. Downstream eligibility comes from a current platform observation, not a completion message.
  7. Keep chat read-only by default. Translate an approved request into a durable scoped job before execution begins.
  8. Show uncertainty honestly. Queued, offline, processing, stale, and needs review are useful states, not design blemishes.

The result is not five autonomous colleagues wandering through websites. It is one controlled publishing operation presented through five understandable responsibilities. That may sound less magical. Magic has historically been weak on audit trails.

The next complication is what happens when an agent may have clicked Submit but lost the response. The safe answer is covered in how to prevent duplicate audiobook submissions after a browser failure.

About the author

Mick Southerland is building Solti's managed bulk audiobook publishing operation for owners of finished book catalogs. These field notes document the engineering choices, platform complications, and operating lessons behind publishing at scale.