Skip to main content

Command Palette

Search for a command to run...

The Retrieval Layer Doesn't Know Which Document Is Current

Updated
•6 min read•View as Markdown

Building a Knowledge Base from Scratch, EP03

When a RAG system gives a wrong answer, the debugging instinct is to blame the retrieval model, the parameters, the reranker. Anything downstream. Here is a post about the opposite case: failures that were guaranteed before a single query ran, in the documents themselves, with request IDs.

The setup continues from EP02, where I built a small knowledge base for a fictional home-goods store and verified that retrieval works: cross-file recall, semantic matching, threshold filtering, all real. This time I polluted it on purpose. Five new documents went into the same library, chosen for how real they feel rather than how clean: an expired 2023 expense policy (nobody deleted it), the 2026 rewrite that contradicts it in six places, the same freight table in two carriers (sectioned Markdown versus a merged-cell Excel with a stale note rotting in sheet two's corner), and a 5,000-word support manual. Eight documents total, 27 slices.

Failure 1: retrieval has no version awareness

The opening query was deliberately mundane: "报销发票怎么提交", how to submit invoices for reimbursement. The top hit, score 0.7226, was the 2023 policy: paper invoices, stapled, to the accountant within 15 working days. The current 2026 policy trailed at 0.7055: e-invoices, photographed, uploaded within 30 days. A 0.017 gap, decided entirely by which document's phrasing better resembles the question.

I ran three queries against the pair. Every single one came back with both versions neck and neck, gaps from 0.017 to 0.05, and the top spot flip-flopping with the wording. The files were literally named "2023旧版" and "2026现行版", expired and current. Filenames don't participate in retrieval. There is no mechanism in this layer that knows time, currency, or supersession. If both versions are in the library, both will be served. Which one a downstream assistant would then repeat is a question this episode didn't run, I stopped at retrieval. The exposure is obvious without running it: the 2023 rule arrived with stapler-level detail, the 2026 rule with a photo upload, and nothing in the library says which of the two is in force. On the one thing I did measure, the stapling ritual beat the app upload.

The engineering conclusion is uncomfortable for people who want to solve this with parameters: this failure mode is not a retrieval problem. No amount of top-k tuning makes an expired document less retrieved when the query wording favors it. The fix is upstream, at ingestion, and it is deletion.

Failure 2: the carrier decides usability, and the score lies

The freight-table pair carried identical information. Asking "新疆买沙发能发货吗", can a sofa ship to Xinjiang, the raw Excel's fragment won at 0.7982:

时效运费:6-8天不发大件(沙发床垫餐桌)差价15

The word "sofa" sits right in the fragment, so the score is honest. But the fragment is unusable: 15 RMB is the difference from what? Is "6-8 days" dispatch or delivery? The clean Markdown's complete rule came second at 0.6894. A lower score, a forwardable answer.

This inverts how most teams audit retrieval quality. Score-first auditing would approve the fragment and reject the rule. Content-first auditing, reading whether the hit can actually answer the question, gives the opposite verdict. Score is a within-query ranking signal. It was never a quality metric, and the moment your source documents degrade into fragments, it becomes an active liar.

The spreadsheet's second crime is inflation. The 6KB file expanded into 11 of the library's 27 slices (the console's slice view counts exactly that), about 40%, and its corner note, "以上如有变动以客服最新答复为准(2024.6 更新)", rode its high-frequency words into the top five of six out of ten queries. One carelessly maintained sheet set the library's noise floor. Format isn't presentation. Format is what the chunker sees, and the chunker is the last thing standing between your spreadsheet and every future query.

The third result was accidental and the best lesson of the batch: the free-shipping threshold query returned 满99包邮 from the new table and 满59包邮 from the EP02 FAQ, both correct when written, never reconciled. Retrieval served both without comment. It doesn't arbitrate. Nobody told it to.

Failure 3: what the chunker rewards

The manual finally forced smart chunking to show its behavior, and three things emerged that builders should internalize.

Titles ride inside slices and participate in matching. The query about sofa sizing hit 0.9108 on a slice whose title chain reads "2.2 尺寸推荐|2.3 比价与砍价|2.4 催券与活动咨询", section headings serving as retrieval road signs. Heading hierarchy in source documents is not typography. It is indexing infrastructure.

Cross-chapter slices exist. One returned slice opens with the tail of section 4.5 and then carries chapter five's opening sections. Semantics survived because the adjacent topics actually belong together. The failure mode to fear is the mid-sentence cut, which the official docs list among the three classic chunking defects, and which well-headed documents make unlikely.

Tables keep content and lose syntax. A two-column Markdown table came back flattened into pipe-joined rows. Downstream parsing that expects table structure will not find it.

And one detail that connects to Failure 1: the manual's first line, "版本:v3.2 | 更新日期:2026 年 6 月", landed in slice one. Retrieval ranking ignores it. The generation-stage model reads it. When two documents fight, the version line in the body text is the only artifact downstream that can tell old from new. Write it there.

The part nobody budgets for

None of the fixes require code: delete expired versions, reconcile one fact to one number, defuse or convert spreadsheets, structure long documents with headings, put version and date in the body. All five happen in a document editor before anything reaches the console. Yet in most RAG project plans, data preparation is a line item at best, because the interesting engineering is downstream and this work looks like clerical chores.

The division of labor matters as much as the list. The person holding the CLI cannot judge whether the 2026 policy supersedes the 2023 one; finance can. The checklist belongs to the document owners, and the verification command belongs to whoever runs the library: pick the rule you trust least, ask about its procedure, and if the top two hits are two versions of the same document, the checklist wasn't finished.

Reproduce it

npm install -g bailian-cli
bl auth login --api-key sk-xxxxx

Build a library with your own documents, but don't clean them first. Add one expired version and one merged-cell spreadsheet on purpose, then run ten queries and watch which passages win. It's the cheapest possible rehearsal of the failure you'd rather discover in an experiment than in a customer's screenshot.


All ten retrievals ran for real on the Bailian CLI against the zj0knmrbye knowledge base; request IDs are kept in the project repo. CLI install: Bailian CLI docs. API key: get one free, new accounts include free quota for 90 days.

More from this blog

Z

zzc

28 posts