A duplicate audiobook submission usually begins with a reasonable instinct: something failed, so try again. In Solti's bulk audiobook publishing operation, we learned that this instinct is safe for some technical failures and dangerous for the one failure that matters most—losing contact immediately after pressing Submit.
Imagine that the worker reaches Spotify’s final review page, confirms the correct title and account, and clicks Submit. The browser connection drops before the confirmation is saved. Did Spotify receive the book? The worker does not know. Pressing Submit again may be harmless, rejected, or duplicate the work. Marking the title failed may send an operator down the wrong path. Calling it complete would be fiction.
The only honest state is uncertain. Once a system can say that word, it can recover safely.
Retry is not the same as recovery.
A retry repeats an operation. Recovery restores reliable state. Those sound similar until the operation changes an external platform that does not share your database transaction.
If a network request fails before it reaches the platform, repeating it may be safe. If an audio upload stops halfway through and the platform exposes no completed asset, restarting may be safe. But if the platform could have accepted the submission, the same replay can create a second record or move the wrong record forward.
Do not convert ambiguity into failure merely because failure is easier to automate. Record the uncertain attempt, inspect the original external record, and decide from evidence.
This is one reason “done” cannot be the only publishing status. A worker can finish its script while the business outcome remains unknown.
Persist the intent before the click.
Before any final external mutation, our system needs a durable statement of intent. That record says which internal book, publishing profile, account, platform lane, external record, and action are involved. It also carries a unique operation key and the asset revision being submitted.
This must be written before the click, not afterward. If the computer stops one millisecond after the platform accepts the title, a replacement worker still needs to know what was attempted. A log message left inside the lost browser session is not durable intent.
The intent record moves through controlled states such as prepared, executing, succeeded, failed, or uncertain. It does not disappear when the worker restarts. The next worker claims the same operation and follows its recovery rule rather than inventing a fresh one.
Use stable platform identity, not titles.
Recovery depends on finding the same external record. Titles and filenames are poor keys: catalogs contain editions, subtitles, translations, and remarkably similar names. The system should preserve the stable platform identifier or review URL as soon as it becomes available and bind it to the internal book ID.
At the final review step, the worker confirms the exact book, account, and platform identifier before publishing. After any interruption, it returns to that identifier. It does not search the catalog and pick the first familiar title.
Account identity matters just as much. An external ID in the wrong publishing account is not the right record. Every recovery check must verify that the current authenticated account matches the book’s publishing profile before changing anything.
Treat a disconnect after Submit as uncertain.
Uncertain is a designed state, not an error message we forgot to rewrite. It blocks automatic replay and records exactly why the outcome cannot yet be classified. For example: the enabled Submit control was activated once, the browser disconnected, and no confirmation or catalog observation was captured.
The worker should save the last known review URL, platform ID, operation key, timestamp, and account binding. It should not click the button again during bounded page recovery, and it should not create a new draft just because the old page is unavailable.
There is no true cross-system “exactly once” transaction between our database and a third-party browser workflow. The platform cannot commit atomically with our ledger. We approximate exactly-once behavior with durable intent, stable identity, one authorized click, and reconciliation before replay.
Reconcile the original record before replay.
Reconciliation asks the platform what happened. Open the correct account, locate the original record by its stable identifier, and inspect its current state. Save the raw observation, the normalized state, and the evidence time against the same internal book.
If the platform shows submitted, processing, accepted, or live, the operation succeeded even if the original worker never recorded its response. Update the ledger from that evidence and continue monitoring. Do not submit again.
If the platform still shows an editable draft and there is reliable evidence that no submission occurred, the original operation may be retried using the same key and record. If the platform cannot be reached or the identity is ambiguous, the book remains uncertain and needs review. Waiting is cheaper than duplicating a catalog.
This recovery works best when all publishing agents share one ledger. A replacement worker can see the prior intent and evidence without depending on the previous agent’s memory.
Retry only failures that are safe to repeat.
Automatic retry should be reserved for transient operations with known outcomes: a timed-out read, a temporary page load problem before any mutation, or a recoverable transport failure where the platform confirms that nothing was created. We use bounded attempts and backoff instead of an infinite loop.
Invalid metadata is not transient. A rejected cover is not fixed by uploading it three more times. An unknown submission outcome is not a failure. Platform processing is not a reason to resubmit. Each of these conditions needs a different next action.
A practical retry policy might allow three attempts for safe transport errors, stop immediately for invalid data, schedule a later observation for processing, and route uncertain mutations to reconciliation. The important part is that the category is determined by evidence, not by how impatient the queue has become.
Let one uncertain book block one book.
At catalog scale, exceptions are normal. One title will have a bad cover, another will be processing, and a third will lose its browser session at precisely the least charming moment. The queue must isolate those conditions by book, lane, profile, and operation.
An uncertain Spotify submission should stop further Spotify mutations for that book. It should not stop a different book whose identity and job are clean, and it should not erase verified Google results. Unique job keys and scoped worker leases keep the blast radius small.
This is where specialist roles help. The Spotify worker owns the reconciliation for its lane while the oversight view exposes the exception. The structure described in our five-agent publishing architecture keeps the work understandable without scattering the facts.
The safe recovery checklist.
When an automated audiobook submission is interrupted, use this order:
- Stop automatic replay. Mark the operation uncertain if an external mutation may have occurred.
- Preserve the original intent. Keep the operation key, book ID, profile, account, platform ID, URL, asset revision, and timestamp.
- Verify the account. Confirm the authenticated platform identity matches the book’s binding.
- Inspect the original record. Reconcile by stable external ID, not by a fuzzy title search.
- Save fresh evidence. Record the raw observation, normalized platform state, and verification time.
- Classify the outcome. Succeeded, safely failed, still processing, or still uncertain are different results.
- Replay only a proven non-event. Reuse the same identity and operation key; never create a fresh submission to escape ambiguity.
- Keep the exception scoped. Block the affected book and lane while unrelated verified work continues.
This design does not eliminate platform failures. It prevents our recovery logic from making them worse. In bulk publishing, that is a meaningful form of progress: the system can be interrupted, restarted, and still know enough not to do the most expensive thing twice.