It was 6 a.m., four days before our request for information (RFI) deadline, when I awoke to a message from my colleague: “Something’s wrong with the demo doc.”
I’d spent the past three weeks treating Claude as my secret weapon. Once a year, analyst firms send software vendors a monster questionnaire to complete including hundreds of questions, plus hours of live product demos, all due in a brutal window.
It’s an enormous amount of work, across an enormous number of contributors, in a painfully short window. I can’t move until product confirms a detail. Product can’t confirm until engineering checks a repo. Someone has to track down that perfect customer example using the specific use case you want to highlight.
At its core, this is a knowledge-retrieval problem. A huge amount of what you need lives nowhere except in people’s heads: the way a specific customer actually uses a feature, the precise technical implementation behind a launch, confirming the roadmap plan.
Why bring up this behind-the-scenes story? It’s the same thing that happens when an intranet page never gets retired after something changes, or when an AI assistant confidently repeats what used to be true. A huge amount of what you need lives nowhere except in people’s heads. And that’s true whether you're prepping an RFI or updating a policy page.
The obvious fix
So this year, I decided Claude was going to be my accelerant. It seemed like the obvious move, a large language model, built to synthesize information quickly, thrown at a mountain of documents and a wall of questions.
It should have been the accelerator I needed, but instead, it gave me answers that contradicted each other. Outdated information dressed up as current fact. When I narrowed its scope to “just these documents,” the answers got cleaner but incomplete.
I also tried telling it which sources to trust over others (“this doc beats that doc”) and that just created a new problem: anything not living in a “top priority” document vanished, even when it mattered. I was cross-checking so much that I couldn’t tell anymore if AI was saving me time or costing me it.
Then came the 6 a.m. message.
Buried in the document we’d built our demonstrations on were inferences — confident, plausible, and wrong — that had quietly rewritten what our product could actually do. Even worse, these weren’t the kind of errors a spot-check would catch. You needed deep, specific product knowledge to even notice something was off.
What followed was four straight 18-hour days. Me and three colleagues, tearing apart content built on a foundation that had drifted from the truth, rebuilding it from scratch before the clock ran out.
We all know AI is only ever as good as the knowledge foundation it draws from. But increasingly I’m learning that “good” isn't just clean. My source documents weren’t sloppy, they were frozen. Someone’s understanding of the product had moved on, and the document never caught up — so that’s exactly what got surfaced back to me.
We caught the problem because a sharp-eyed colleague happened to notice, days before it would've been too late. That’s not a system. . . that’s luck. It makes me wonder if the real opportunity isn’t asking AI to flag what's already outdated, but something harder: Could it help surface tribal knowledge as it forms, instead of waiting for someone to write it down and finding out months later it was never captured?