All posts
2 min readby Romiel Inolino

Your RAG demo took a weekend. Making it reliable enough to run the business is the real project

RAGAI AutomationKnowledge BaseAI AgentsOperations

Pointing a chatbot at your Google Drive is easy. Getting answers your team can trust every day is a different job.

What happened

A VentureBeat article by Abdullah Sayyad, published September 27, argues that a RAG demo is quick to build but real company data is where "the real engineering work begins." Documents disagree, some are outdated, useful facts sit in spreadsheets and PDFs, different employees have different permissions, and search gets worse as the knowledge base grows.

The piece makes several practical points:

  • Data quality beats model choice. A better embedding model cannot fix a knowledge base where nobody has decided which documents are authoritative.
  • Chunking is an architecture decision. Manuals should keep their structure, while policies need metadata like department, region and effective date.
  • Vector search alone misses exact matches, so hybrid retrieval that mixes semantic and keyword search matters for things like order numbers and SKUs.
  • Test retrieval and generation separately, or you cannot tell which part failed.
  • Access control has to happen before restricted content reaches the model's context.

It cites Ring, Amazon's home security company. Per the AWS write up, Ring tagged support content by locale, served 10 international regions from one central knowledge base, split content into ingestion, evaluation and promotion workflows, and reduced the cost of scaling to each additional locale by 21%.

My take

Many "our AI bot gives wrong answers" complaints are not model problems. They are content problems: three versions of the refund policy, a price list from last year, and a folder everyone can read even though half of it is HR.

When I scope a knowledge assistant for a team, the first week is boring on purpose:

  1. Pick the sources that count as the truth and archive the rest.
  2. Tag documents with owner, date and audience so retrieval can filter.
  3. Write 30 to 50 real questions from staff and check that the right passage comes back before any answer is generated.
  4. Set a refresh schedule so the index does not quietly go stale.

None of this is exciting, but it is the difference between a demo people clap for and a tool people use on Monday. If your team already has a bot that nobody trusts, the fix is usually in the data layer, not a model upgrade.

Sources are listed below.

More posts