A publishing review shows how I keep AI answers tied to sources, human decisions, and work that can move between tools.
In August, I directed an AI-assisted review of 287 authors in our publishing production records. For each author, we needed to identify the right book and examine the public evidence. The record had to remain understandable after the agent's conversation ended, because another person might need to correct it months later.
Our first searches sometimes used an author's name without a title. The local file scan also skipped EPUBs and stopped before it reached manuscripts stored in deeper folders. For some authors, an EPUB was the only title file in the folder. I corrected the search method as the gaps became clear.
One error needed knowledge of the production work itself. The scan found a biography and a book description, then treated their presence as evidence that the manuscript was missing. Those documents had been extracted from a manuscript. I directed another search of the production files, including the nested folders where our team stored manuscripts. A stronger model would still have received the same incomplete scan.
Agents gathered evidence, checked the author and book, and challenged claims before returning rows for the worklist. One coordinating process wrote the shared file to reduce overwrite risks. It still made a mistaken merge when fuzzy matching joined records that belonged to different people. I repaired the merge. A future version should use a stable record identifier and stop for review when identity is uncertain, and the source would still need its own check.
The worklist also had to express uncertainty honestly. If a search did not find evidence for an award, the award remained unconfirmed. A public page that repeated an author's own claim needed to remain distinguishable from independent confirmation. Those distinctions affect what a colleague can safely write about the book.
In Obsidian, I keep source material distinguishable from a working interpretation and an accepted decision. An unresolved question stays visible rather than being smoothed into a fluent answer. I want to know which file led to a conclusion, what was corrected, and whether someone has accepted the result for a particular use. Eleven years running a team of auditors, alongside my publishing work, taught me to expect that a later reviewer will ask precisely those questions.
The mistaken merge showed why a write restriction and an identity check answer different questions. If I give an assistant permission to take a more consequential action, the service making the change needs to check the exact approval and the record it applies to. A sentence from the model saying it has permission is only another claim to examine. The publishing review did not demonstrate a complete approval system, it clarified the need why I would require one.
Before assigning the narrower search step to a smaller model, I directed a comparison of two model tiers on 15 difficult author lookups with the same prompts. The recorded review found no fabricated fields in that sample on either tier. It supported using the smaller model for that search step, with checks continuing during the wider review. Fifteen cases were useful for one decision.
The review used tiers from one provider. I would test a replacement against the cases that caused trouble: similar names, missing local files, and claims with weak public evidence.
My local AI gateway connects several local tools and explicitly chosen remote services through a common interface. It helps me change the service used for a task without rebuilding every connection. I keep the sources and decisions in readable files that survive the end of an account or subscription. Any replacement still needs to respect what each field meant when it was created by the previous writers.
For a person or a small business choosing AI, I would begin with one piece of work they need to finish. How much time does the assistant save after they check the result? Who backs up the files and repairs a broken connection? A lower model bill can still leave someone with more work if every answer needs correction. A local substitute can also create dependence on its builder, and additional cost through time spent to maintain the local system.
ProseGuard, a local writing checker I built, illustrates a more limited role for software. It flags repeated phrases and punctuation habits against clear rules that can be written for one person's voice. It does not decide whether a paragraph says what its author means. I use it alongside my voice profile and a human reading of the text. A zero-finding report is useful, but it cannot supply an opinion, verify a production record, or show that an interface feature works.
My LCARS project is a hobby inspired by the Computer in Star Trek: The Next Generation. I am developing a visual and voice interface around my knowledge work. I want the screen to show where an answer came from and whether it is a candidate, a checked record, or an open question. The project remains a prototype for personal use.
My Master of Fine Arts in Sculpture taught me to consider what a person can perceive. A voice assistant can sound sure even when its answer is provisional. The screen should give someone a chance to inspect the source before acting on it.
The next useful test is a handover. Can someone else resume the work or review from its sources, corrected matches, and open questions without relying on my memory or the original model? That would show whether the knowledge remains usable after either the model or its builder is gone. Needles to say, there is already a tested and proven solution exactly for that in my Knowledge Management System.
Please sign in to leave a comment.
No comments yet. Be the first to comment!