Whether an internal knowledge assistant gives correct answers depends less on the model's name than on decisions made before any code is written: which documents enter the index, how they are split, how permissions travel with them and who catches a wrong answer. This guide walks through those decisions in the order a business owner and an IT lead usually face them in RAG development.
The page above describes what the assistant does. Here we focus on how, and on when not to build one at all.
01Five questions that test your readiness
Before you budget for RAG development, you can check with five questions whether your document archive can carry an assistant. Answer honestly; if more than two answers are no, documentation work is the smarter first investment.
- Repeated questions: Do the same people get asked for the same information several times a week? Then the saved time is real; if every question needs judgment, the gain stays small.
- Written answers: Is the answer to common questions in a policy, manual or record? Knowledge that lives only in someone's head cannot be indexed.
- Clear current version: Can an employee tell within minutes which copy of a document is in force? If they cannot, the assistant cannot either.
- Named owners: Is a person or team responsible for keeping each document group current?
- Clean permissions: Do shared drive permissions reflect real roles, or have folders been opened to everyone over the years?
The last point is often skipped. The assistant inherits source permissions as they are, so a drive where everyone sees everything becomes a visible problem once people can query it in plain language. Even a negative result is useful: the inventory, version cleanup and ownership table it triggers support any later assistant.
02Pick the first use case by company type
Start where repeated questions pile up and where a wrong answer can be corrected without lasting harm. The same foundation is fed with different documents and needs different limits depending on the business.
- Manufacturer: Maintenance instructions, quality procedures and past nonconformance reports; a shift lead sees how a fault was fixed before. Safety steps always link to the original document.
- Law firm or consultancy: Contract templates, precedent filings and internal memos; the assistant finds the right template and clause, and the lawyer makes the legal call. Separation between client matters is decisive here.
- Software company: Technical docs, architecture decision records and resolved tickets; new engineers learn why a service works the way it does.
- Franchise or multi location retail: Promotion rules, return procedures and store operations manuals; staff in a branch see the current rule without calling head office.
- Insurance broker or financial services firm: Policy wordings, product guides and process notes; an agent finds the relevant coverage clause before replying and still owns the interpretation.
Choose with two criteria: is question volume high, and can a wrong answer be reversed? An expense rule mistake is easy to fix; a safety step mistake may not be. For high risk groups, let the pilot show the relevant passage instead of a written summary.
03The path a question takes through the assistant
Between the moment a user types a question and the moment an answer appears, seven steps run. Because each step can be tested on its own, errors can be traced step by step too.
- Query preparation: Abbreviations and internal jargon are expanded, and vague phrases such as "last year's policy" become a date range.
- Hybrid search: Vector search (turning text into numeric representations of meaning and measuring similarity) catches synonyms; keyword search does not miss exact strings such as part numbers or clause numbers.
- Reranking: A second model reorders candidate passages by relevance, filtering out noise from the first pass.
- Permission filter: Every passage the user may not see is removed before the model is called; a model cannot leak text it never received.
- Context assembly: Remaining passages reach the model with heading, date and file name; if two versions cover the same topic, the newer one is marked.
- Answer and citation: The model writes only from the supplied passages and ties each claim to a source; when support is thin, it says so.
- Logging: The question, passages used, model and index version are stored, and the user can flag the answer as wrong or incomplete.
Ask every vendor how each step is tested; grading only the final answer cannot show whether a fault sits in retrieval or in writing.
04How you split documents shapes answer quality
Cutting documents into fixed length pieces is a common source of wrong answers; chunking should follow the document's own structure. If clause 4.2 of a policy lands in one chunk and the paragraph describing its exception lands in the next, the model may state the rule without the exception.
We therefore split along heading hierarchy, clause numbers and paragraph boundaries rather than character counts, and keep these details intact:
- Heading chain: Each chunk carries its parent headings, so a short paragraph titled "Scope" still knows which policy it belongs to.
- Tables: Rows are stored with the header row; otherwise a value loses the column that gives it meaning.
- Definitions: A definitions section is linked to the chunks that use those terms and retrieved with them when needed.
- Metadata: Document type, effective date, owner, language and confidentiality level are written onto every chunk; filters work with these labels.
- Scanned pages: Text recognition output is spot checked before indexing; a misread number produces a wrong answer that looks right.
Extracting fields from structured paperwork such as invoices or application forms is a different job; see AI document processing for that. A knowledge assistant aims to make text searchable without breaking its meaning.
Duplicates distort results as well: three copies of one file fill three slots in the result list and crowd out real variety. Comparing file hashes with a hash generator lets you remove exact copies before indexing.
05Connecting source systems without blind spots
For every source system, three things are solved separately: reading content, catching changes and carrying permissions over. Miss one and the assistant either goes stale or sees more than it should.
- Shared drives and cloud storage: SharePoint, OneDrive and Google Drive expose change information and access lists; the connection runs through a service account with read access only.
- Wikis and intranets: In Confluence or Notion, page hierarchy becomes section headings and archived pages are dropped.
- Helpdesks: Only resolved, approved tickets from tools such as Zendesk are pulled in, with customer names and contact details masked first.
- Email archives: Often the riskiest source; only specific shared mailboxes are added, with a stated purpose and a time limit.
- Business applications: Records in an ERP or a CRM such as Salesforce are live data, not documents; querying them narrowly at question time is sounder than copying them into the index.
Deleted files are often forgotten. A document removed or restricted at the source must leave the index in the same sync cycle, or the assistant keeps quoting something nobody can open; we test this on purpose during setup.
If the assistant has to live inside an internal application rather than Teams or Slack, that interface is planned through custom software development with your sign in model.
06Choose model and hosting by document sensitivity
Model choice is as much about where data goes as about answer quality, and it follows the sensitivity of each document group.
Two kinds of model run in an assistant. An embedding model (one that turns text into a list of numbers representing its meaning) builds the index; a language model writes the answer. If your documents are in several languages, test the embedding model in each of them, since a model that shines on English benchmarks can miss relevant passages elsewhere.
- Cloud API: Strong answer quality and low maintenance; only passages relevant to each question are sent. If those passages contain personal data, transfer rules apply.
- Cloud with regional processing: Some providers keep processing in a chosen region, which simplifies contract review but does not settle the legal question alone.
- Open model on your own server: Documents never leave your infrastructure; hardware, updates and quality tracking become your responsibility.
- Hybrid routing: General policies go to a cloud model while HR and legal files go to a local model, routed automatically by confidentiality label.
Our page on local LLM deployment covers the hardware and operations side and is worth reading before you decide.
One detail to know early: switching the language model leaves the index untouched, but switching the embedding model means rebuilding the index. In RAG development we therefore compare embedding models on your own documents during the pilot and treat any later change as a planned release.
07Human review, feedback and logging
A knowledge assistant retrieves information rather than making decisions, yet you should define in writing, before launch, which topics get a referral instead of an answer, where feedback goes and how long logs are kept.
Most internal questions can be answered directly. For individual pay, disciplinary or medical matters, and for legal commitments that will reach a customer, the assistant should show the relevant document and point to the responsible team instead.
- Feedback routing: A flagged answer goes to the owner of the cited document, who decides whether the document or the retrieval setting needs fixing.
- Weekly review: During the pilot, flagged answers and unanswered questions are read once a week with document owners and the project lead.
- Change approval: No prompt, index setting or model change goes live before passing the question set, and who signs off is agreed in advance.
- Log scope: Question text, passages used, answer, version data and feedback are stored, with log access limited to a few roles.
Set retention to balance useful history against keeping employee data no longer than needed. Stating in writing that logs will not rate individuals keeps usage honest; people do not ask real questions of a tool they think is watching them.
08GDPR, employee data and the EU AI Act
Data protection touches an internal assistant in two places: personal data inside indexed documents, and the question logs of the employees who use it. Assess the two separately.
- Personal data in documents: HR files, customer correspondence and contracts contain personal data. Ask whether each document group is truly needed; not connecting an unnecessary group is the strongest safeguard.
- Processors: When documents with personal data reach a model provider, that provider acts as a processor and needs a data processing agreement under Article 28 of the GDPR.
- International transfers: If data leaves the EU or EEA, Chapter V of the GDPR requires an adequacy decision or appropriate safeguards such as standard contractual clauses. UK and US based teams should check their own transfer and privacy rules with counsel.
- Sensitive groups: Health records and similar files are either kept out of the index or processed only by a local model with narrow access.
Article 4 of the EU AI Act, as amended in 2026, asks organizations that deploy AI systems to take measures to support their staff's AI literacy. We treat this less as a compliance box and more as a precondition for using the assistant well; corporate AI training can close that gap for your teams.
This is not legal advice; finalize the transfer basis and employee notice with your legal adviser.
09Run the pilot with written criteria
A pilot runs with one team, a limited document set and success criteria written down beforehand. Without criteria, nobody can make a decision when the pilot ends.
- Write the scope: Which team, which document groups and which question types, plus a list of what is out of scope.
- Collect the question set: Gather real questions and correct answers with their source passage from the team, mixing easy, hard and not in the documents cases.
- Measure the baseline: Time sample lookups today; a working hours calculator turns that into weekly team hours.
- Finish internal testing: Run the question set and test citations, the permission filter and the deleted document scenario separately.
- Open to part of the team: A subset starts using it, and feedback plus unanswered questions are reviewed weekly.
- Hold the decision meeting: If criteria are met, add a second team or document group; if not, use the logs to locate the problem in documents, retrieval or the model.
Tie pilot length to question volume rather than the calendar; a pilot few people used is simply inconclusive. When you expand, each new document group arrives with its own questions, so older groups are retested every time.
10Five indicators that show real value
Value is tracked with indicators that separately measure whether retrieval finds the right passage, whether answers stay faithful to sources and whether the team actually adopts the tool, not with a single satisfaction score.
- Retrieval hit rate: For each question in the set, is the correct passage among the top results? If this is low, the problem lies in chunking and search, not the model.
- Faithfulness: Is every claim in the answer actually written in the cited passage? Measured by human review on a sample.
- Correct "I don't know": For questions with no answer in the documents, how often does the assistant say so instead of inventing one?
- Return usage: Do pilot users come back after the first week? This separates curiosity from habit.
- Questions to experts: Are routine questions to senior staff declining, tracked with a short survey or by counting messages in an internal support channel?
Recalculate these with the same question set whenever the index or model changes, since one measurement cannot reveal slow decay. If you want a financial figure from RAG development, base it on your own measured search time and headcount, not on ratios borrowed from other companies.
11Where retrieval augmented generation falls short
RAG is strong at finding and relaying what documents say, and weak at inferring what they do not say, counting across an archive and resolving contradictions. Knowing this upfront sets the right expectations.
- Aggregate questions: "How many contracts this year include a penalty clause" requires scanning the whole archive; RAG retrieves only a handful of relevant passages and may produce an incomplete number. Reporting tools handle such needs.
- Conflicting documents: When two valid documents disagree, the assistant should surface the conflict rather than pick one; the lasting fix belongs to the document owner.
- Fluent but wrong: A model can build a convincing sentence on an incomplete passage. Citations reduce this risk without removing it, so users should open the source on critical topics.
- Hidden instructions: Text planted in an external file may try to steer the model; read access only and no outbound actions are essential for that reason.
- Stale permissions: If someone changed roles and their old access was never removed at the source, the assistant keeps honoring it.
Overreliance is the quietest risk: the more often the assistant is right, the less people check sources. Handover training therefore shows which question types it handles poorly.
12Common mistakes and what to do instead
Most RAG development mistakes are about scope and order, not technology. Each of these six can be prevented before setup.
- Connecting the whole drive at once: Drafts, copies and old versions enter the index too. Start with one document group and add others only after cleaning them.
- Judging by demo questions: A vendor's own questions always look good. Insist on a set built from your team's real questions, including ones the documents cannot answer.
- Adding permissions later: Retrofitting a permission filter may force an index rebuild. Make access control part of the index from day one.
- No document owners: If nobody is responsible for updates, the assistant serves outdated facts within months. Put a name on every source group.
- Ignoring shadow AI use: Staff keep pasting company files into personal chat accounts. When you launch, publish a short usage rule explaining why that should stop.
- Measuring only satisfaction: Thumbs up ratings hide wrong answers. Track question set results, faithfulness and the unanswered list together.
What these share is treating the assistant as a one time install rather than a system that changes with your documents.
13Questions to ask a RAG partner
When comparing RAG development proposals, look past the demo to how the team catches errors, carries permissions and what it leaves you at handover. Ask for written answers to these:
- Permissions: How and how quickly does a permission change at the source reach the index, and is a deleted document test part of acceptance?
- Testing: Whose questions form the question set, are retrieval and answers measured separately, and will you see results after every update?
- Data path: Which passages go to which provider and region, and can they show documentation that the provider does not train on API data?
- Ownership: In whose name are accounts, index settings, prompts and the question set; could another team take the system over?
- Maintenance: How are new document types, model changes and quality drops handled after launch?
Asking for references is fair. We have not yet delivered a knowledge assistant for a client we can name; our automation, software and web work is on our references page. So we present our method in writing with pilot criteria and suggest you judge it with your own questions. Build, connectors and upkeep run within our AI automation services; a proposal that leaves maintenance vague should weigh heavily in your comparison.
14Making the decision and the first step
If you answered yes to at least three readiness questions and one team clearly carries a load of repeated questions, you are in a good position for a pilot; if not, document cleanup comes first. To make the first conversation productive, gather:
- The team losing the most time and three to five typical questions from it
- The systems and document types where those answers live
- Document groups you treat as sensitive and whether they may leave your infrastructure
- Where people will use the assistant: browser, Teams, Slack or an internal portal
- The criterion that would make the pilot a success for you
With this, a short call is enough to agree on pilot scope, the data path and high risk document groups for your RAG development project. You can request that call through our contact page.
Cost depends on document volume, number of source systems, permission complexity and hosting choice. Fixed scope options for discovery and the first workflow are in the AI automation pricing section, and a written quote follows once scope is clear. If the readiness test says not yet, we tell you plainly, and the outcome of the call becomes a cleanup list for your documentation.