AI solution

RAG Development for Internal Knowledge

A RAG knowledge assistant gives your team the answer that is buried somewhere in company documents, together with the exact file and section it came from. We do not treat it as a generic chat tool but as internal infrastructure, built with its document sources, access permissions, update cycle and measurement.

Answers with citationsPermission aware retrievalAutomatic document updatesWorks in your internal toolsUnanswered question report
  • Google Partner
  • Talha Aslan and team
  • English, German, Turkish

In short

RAG development means building an assistant that uses retrieval augmented generation: when an employee asks a question, the system first finds the relevant passages in your company documents, then a language model writes an answer based only on those passages and cites them. Each user only reaches documents they are allowed to open, the assistant says so when it finds no support, and its index updates as documents change.

Talha Aslan and teamLast updated:

When you need one

Where does company knowledge get lost?

Not every company needs a knowledge assistant. If you have few documents, if knowledge lives in a handful of people's heads and changes weekly, or if your documents are out of date, tidying up the documentation is the better first investment. If the situations below sound familiar, an assistant is worth considering.

One answer means searching folders and chasing colleagues

Leave policy, purchase approval limits, a product's technical tolerance: the answer is written down somewhere, but nobody remembers which drive or which version. The employee ends up messaging someone who knows, and two people lose focus instead of one.

New hires take a long time to get up to speed

Onboarding documents exist, yet nobody reads them end to end. Newcomers spend their first weeks asking senior staff the same questions, and the people passing on knowledge fall behind on their own work.

Company files get pasted into personal AI accounts

To save time, staff copy contracts or customer files into chat tools on their private accounts. Where that data goes, who can see it and whether it is used for training all remain unclear.

An outdated version is taken as current

Several versions of the same procedure sit in different folders. Search may surface the oldest one, and someone acts on a rule that no longer applies.

Our approach

An assistant grounded in documents, aware of permissions, kept current

We start from a document inventory, not from software. Together with you we map which knowledge sits in which system (shared drives, intranet, wiki, support tickets, PDF archives), which documents are current and approved, and who may access what. Separating old versions and drafts at this stage takes far less effort than chasing wrong answers later.

We split documents into meaningful passages and write each one into a search index with its heading, date and access label. When a question arrives we run semantic search and keyword search together, filter the best passages by the user's permissions and pass only those to the language model. Build, source connectors and upkeep run as part of our AI automation services.

This assistant is meant for internal use; if you are looking for a chat assistant open to your customers, see our AI chatbot development page. If the assistant has to live inside your own portal or business application, the interface is built through custom software development.

  • Answers only from retrieved document passages
  • File name and section link with every answer
  • Users only see documents they may access
  • Index refreshes when documents change
  • Unanswered questions reported as content gaps
Anatomy of a RAG knowledge assistant
  1. Document sourcesDrives, wiki, intranet and PDF archive
  2. Chunking and indexWith heading, date and access label
  3. Hybrid searchMeaning and keywords combined
  4. Permission filterPassages the user cannot see never reach the model
  5. Cited answerQuote, file name and link
  6. FeedbackWrong or incomplete answers get flagged

Answer quality depends on how well the documents are organized as much as on the model; the index, permission and feedback layers are part of every build.

Which assistant?

Who the assistant serves shapes how it is built

The same foundation is set up differently for each team, so we first pick the knowledge area where most time is lost.

HR and administration

Policy and procedure assistant

Answers questions about leave, expenses, purchasing and health and safety rules by pointing to the relevant clause.

  • Only versions currently in force are indexed
  • Personal matters routed to HR
  • Onboarding question set for new hires

Technical teams

Technical documentation and maintenance assistant

Searches product manuals, maintenance instructions and past fault records and shows the steps with their source.

  • Parsing of PDFs with tables and diagrams
  • Exact matching on model and part numbers
  • Mobile interface for field staff

Sales and support teams

Product knowledge assistant for agents

Finds the answer to a customer question in product documents and resolved tickets; the agent approves any text that goes to the customer.

  • Read only connection to the helpdesk
  • Draft reply ready, final call with the agent
  • Monthly list of recurring questions

Essentials

The building blocks of a trustworthy knowledge assistant

These points matter as much as answer accuracy: they keep the wrong person from seeing the wrong document and make sure mistakes are caught in time.

Document level permissions

The assistant carries access rights from the source systems into the index, so a user cannot read the content of a document they are not allowed to open, not even inside an answer. The OWASP list for LLM applications treats permission aware vector stores as a risk area of its own.

Protection against instructions in documents

An instruction hidden inside a document may try to change how the model behaves. We treat document text as data, never as commands, give the assistant read access only and define no action that sends data outside.

AI literacy

Article 4 of the EU AI Act, as amended in 2026, asks deployers of AI systems to take measures to support their staff's AI literacy. At handover we run a short training session and provide a guide that explains what the assistant does not know; we do this for every team, inside the EU or not.

Processors and transfers

If documents contain personal data, the model provider acts as a processor and signs a data processing agreement under Article 28 of the GDPR. Where data leaves the EU or EEA, Chapter V requires an adequacy decision or safeguards such as standard contractual clauses. Your legal adviser makes the final call.

Tiers that do not train on your data

We send document passages only to commercial API tiers. Anthropic, for example, states that by default it does not use inputs or outputs from its commercial products, including the API, to train its models. For highly sensitive archives, running the model on your own server is also an option.

Version log and rollback

The index version, prompt and model are logged with every answer. Changes are tested against the question set first; if quality drops, we return to the previous index and settings.

Sources: OWASP Top 10 for LLM Applications 2025, LLM08: Vector and Embedding Weaknesses · EU AI Act (Regulation 2024/1689), Article 4 as amended by 2026/1744, EUR-Lex · General Data Protection Regulation (2016/679), EUR-Lex · Anthropic Privacy Center: is my data used for model training? · Lewis et al. (2020), Retrieval Augmented Generation for Knowledge Intensive NLP Tasks, arXiv

Comparison

Generic AI tool or a RAG knowledge assistant?

TopicUploading files to a generic chat toolRAG knowledge assistant
Knowledge sourceWhatever file the user uploads at that momentYour full archive of approved documents
FreshnessFiles are uploaded again every timeIndex refreshes when a document changes
PermissionsWhoever uploads can share anythingSource system permissions are kept
CitationsOften missing or vagueFile name, section and link
Data controlScattered across personal accountsIn the company account, logged and limited
MeasurementNot possibleUnanswered questions, feedback and usage report

Quick check

RAG knowledge assistant feature list

Must haves: is your process ready?

0 of 6 in place Tick the boxes to see how ready you are for automation.

Added as needed

  • Use from Teams or Slack
  • Single sign on (SSO)
  • Text recognition for scanned files (OCR)
  • Multilingual documents and questions
  • Model running on your own server
  • Documentation plan from unanswered questions

We choose which of these you need together during the first call.

Let us choose the documents your assistant should read

Tell us what your teams search for most and which systems hold those documents; we will outline the pilot scope, access rules and a written quote.

Process

From discovery to launch in four steps

  1. First call and discovery

    We listen to your processes in a free 15-minute call. Then discovery maps your tools and tasks, scores the opportunities and ends with a written scope and fee for your approval.

  2. Build and test

    We build the first workflow in your accounts and test it with real but masked examples. Approval steps, error scenarios and alerts go in before anything reaches a customer.

  3. Go live and tune

    We switch the workflow on step by step, watch the logs and adjust thresholds with your team. You get documentation and a short training session.

  4. Monitor and expand

    On the monthly plan, we monitor running workflows, adapt them to model and API changes and add new workflows from the priority list, with a monthly report.

Free tools

Prepare your documents with free tools

Check how readable and long your document texts are, spot duplicate files by their hash, estimate the working hours lost to searching and the rates that matter, and create strong passwords for service accounts.

Content

Readability Checker

Readability score with Flesch (EN), Ateşman (TR) and Flesch-Amstad (DE).

Content

Word & Character Counter

Words, characters, sentences + live checks against Google, Instagram, X limits.

Security

Hash Generator

Create MD5, SHA-1, SHA-256 and SHA-512 hashes of text and files in your browser, verify a download against its checksum and compare hashes. Nothing is uploaded.

Work

Working Hours Calculator

Calculate daily and weekly working hours after breaks, in hours and decimals, and check legal breaks and rest periods for the UK and EU.

Calculator

Percentage Calculator

Percent of a number, what-percent ratio and percent change (increase/decrease).

Security

Password Generator

Cryptographically random strong passwords + strength meter + crack time.

All free tools

How we work

We launch the assistant with one team and a limited document set

We have not yet built a knowledge assistant for a client whose name we can share. So this page describes method rather than results; our team's automation, software and web work is on the references page.

Clean documents first

Old versions, drafts and duplicates are set aside before indexing, and each source gets a named person responsible for keeping it current.

Pilot with one department

The assistant opens for a single team with a limited document set and grows step by step based on the question set and feedback.

Provider independent design

The index and document pipeline are not tied to one model. Our own website's AI powered tools also switch to the next model when one does not respond.

Full handover

Accounts are opened in your company's name; index settings, prompts, the question set and an operating guide are handed over to you.

All references

FAQ

Questions about RAG development

If your question is not here, write to us; we will send you an answer and a written quote.

Next step

Let us scope your first knowledge assistant

Tell us which team loses time searching which documents; after a free 15 minute call we will send the pilot scope and a written quote.

In-depth guide

RAG Development: Readiness, Architecture, Security and Measurement

Talha Aslan and teamLast updated: 15 min read

Whether an internal knowledge assistant gives correct answers depends less on the model's name than on decisions made before any code is written: which documents enter the index, how they are split, how permissions travel with them and who catches a wrong answer. This guide walks through those decisions in the order a business owner and an IT lead usually face them in RAG development.

The page above describes what the assistant does. Here we focus on how, and on when not to build one at all.

Five questions that test your readiness

Before you budget for RAG development, you can check with five questions whether your document archive can carry an assistant. Answer honestly; if more than two answers are no, documentation work is the smarter first investment.

  • Repeated questions: Do the same people get asked for the same information several times a week? Then the saved time is real; if every question needs judgment, the gain stays small.
  • Written answers: Is the answer to common questions in a policy, manual or record? Knowledge that lives only in someone's head cannot be indexed.
  • Clear current version: Can an employee tell within minutes which copy of a document is in force? If they cannot, the assistant cannot either.
  • Named owners: Is a person or team responsible for keeping each document group current?
  • Clean permissions: Do shared drive permissions reflect real roles, or have folders been opened to everyone over the years?

The last point is often skipped. The assistant inherits source permissions as they are, so a drive where everyone sees everything becomes a visible problem once people can query it in plain language. Even a negative result is useful: the inventory, version cleanup and ownership table it triggers support any later assistant.

Pick the first use case by company type

Start where repeated questions pile up and where a wrong answer can be corrected without lasting harm. The same foundation is fed with different documents and needs different limits depending on the business.

  • Manufacturer: Maintenance instructions, quality procedures and past nonconformance reports; a shift lead sees how a fault was fixed before. Safety steps always link to the original document.
  • Law firm or consultancy: Contract templates, precedent filings and internal memos; the assistant finds the right template and clause, and the lawyer makes the legal call. Separation between client matters is decisive here.
  • Software company: Technical docs, architecture decision records and resolved tickets; new engineers learn why a service works the way it does.
  • Franchise or multi location retail: Promotion rules, return procedures and store operations manuals; staff in a branch see the current rule without calling head office.
  • Insurance broker or financial services firm: Policy wordings, product guides and process notes; an agent finds the relevant coverage clause before replying and still owns the interpretation.

Choose with two criteria: is question volume high, and can a wrong answer be reversed? An expense rule mistake is easy to fix; a safety step mistake may not be. For high risk groups, let the pilot show the relevant passage instead of a written summary.

The path a question takes through the assistant

Between the moment a user types a question and the moment an answer appears, seven steps run. Because each step can be tested on its own, errors can be traced step by step too.

  1. Query preparation: Abbreviations and internal jargon are expanded, and vague phrases such as "last year's policy" become a date range.
  2. Hybrid search: Vector search (turning text into numeric representations of meaning and measuring similarity) catches synonyms; keyword search does not miss exact strings such as part numbers or clause numbers.
  3. Reranking: A second model reorders candidate passages by relevance, filtering out noise from the first pass.
  4. Permission filter: Every passage the user may not see is removed before the model is called; a model cannot leak text it never received.
  5. Context assembly: Remaining passages reach the model with heading, date and file name; if two versions cover the same topic, the newer one is marked.
  6. Answer and citation: The model writes only from the supplied passages and ties each claim to a source; when support is thin, it says so.
  7. Logging: The question, passages used, model and index version are stored, and the user can flag the answer as wrong or incomplete.

Ask every vendor how each step is tested; grading only the final answer cannot show whether a fault sits in retrieval or in writing.

How you split documents shapes answer quality

Cutting documents into fixed length pieces is a common source of wrong answers; chunking should follow the document's own structure. If clause 4.2 of a policy lands in one chunk and the paragraph describing its exception lands in the next, the model may state the rule without the exception.

We therefore split along heading hierarchy, clause numbers and paragraph boundaries rather than character counts, and keep these details intact:

  • Heading chain: Each chunk carries its parent headings, so a short paragraph titled "Scope" still knows which policy it belongs to.
  • Tables: Rows are stored with the header row; otherwise a value loses the column that gives it meaning.
  • Definitions: A definitions section is linked to the chunks that use those terms and retrieved with them when needed.
  • Metadata: Document type, effective date, owner, language and confidentiality level are written onto every chunk; filters work with these labels.
  • Scanned pages: Text recognition output is spot checked before indexing; a misread number produces a wrong answer that looks right.

Extracting fields from structured paperwork such as invoices or application forms is a different job; see AI document processing for that. A knowledge assistant aims to make text searchable without breaking its meaning.

Duplicates distort results as well: three copies of one file fill three slots in the result list and crowd out real variety. Comparing file hashes with a hash generator lets you remove exact copies before indexing.

Connecting source systems without blind spots

For every source system, three things are solved separately: reading content, catching changes and carrying permissions over. Miss one and the assistant either goes stale or sees more than it should.

  • Shared drives and cloud storage: SharePoint, OneDrive and Google Drive expose change information and access lists; the connection runs through a service account with read access only.
  • Wikis and intranets: In Confluence or Notion, page hierarchy becomes section headings and archived pages are dropped.
  • Helpdesks: Only resolved, approved tickets from tools such as Zendesk are pulled in, with customer names and contact details masked first.
  • Email archives: Often the riskiest source; only specific shared mailboxes are added, with a stated purpose and a time limit.
  • Business applications: Records in an ERP or a CRM such as Salesforce are live data, not documents; querying them narrowly at question time is sounder than copying them into the index.

Deleted files are often forgotten. A document removed or restricted at the source must leave the index in the same sync cycle, or the assistant keeps quoting something nobody can open; we test this on purpose during setup.

If the assistant has to live inside an internal application rather than Teams or Slack, that interface is planned through custom software development with your sign in model.

Choose model and hosting by document sensitivity

Model choice is as much about where data goes as about answer quality, and it follows the sensitivity of each document group.

Two kinds of model run in an assistant. An embedding model (one that turns text into a list of numbers representing its meaning) builds the index; a language model writes the answer. If your documents are in several languages, test the embedding model in each of them, since a model that shines on English benchmarks can miss relevant passages elsewhere.

  • Cloud API: Strong answer quality and low maintenance; only passages relevant to each question are sent. If those passages contain personal data, transfer rules apply.
  • Cloud with regional processing: Some providers keep processing in a chosen region, which simplifies contract review but does not settle the legal question alone.
  • Open model on your own server: Documents never leave your infrastructure; hardware, updates and quality tracking become your responsibility.
  • Hybrid routing: General policies go to a cloud model while HR and legal files go to a local model, routed automatically by confidentiality label.

Our page on local LLM deployment covers the hardware and operations side and is worth reading before you decide.

One detail to know early: switching the language model leaves the index untouched, but switching the embedding model means rebuilding the index. In RAG development we therefore compare embedding models on your own documents during the pilot and treat any later change as a planned release.

Human review, feedback and logging

A knowledge assistant retrieves information rather than making decisions, yet you should define in writing, before launch, which topics get a referral instead of an answer, where feedback goes and how long logs are kept.

Most internal questions can be answered directly. For individual pay, disciplinary or medical matters, and for legal commitments that will reach a customer, the assistant should show the relevant document and point to the responsible team instead.

  • Feedback routing: A flagged answer goes to the owner of the cited document, who decides whether the document or the retrieval setting needs fixing.
  • Weekly review: During the pilot, flagged answers and unanswered questions are read once a week with document owners and the project lead.
  • Change approval: No prompt, index setting or model change goes live before passing the question set, and who signs off is agreed in advance.
  • Log scope: Question text, passages used, answer, version data and feedback are stored, with log access limited to a few roles.

Set retention to balance useful history against keeping employee data no longer than needed. Stating in writing that logs will not rate individuals keeps usage honest; people do not ask real questions of a tool they think is watching them.

GDPR, employee data and the EU AI Act

Data protection touches an internal assistant in two places: personal data inside indexed documents, and the question logs of the employees who use it. Assess the two separately.

  • Personal data in documents: HR files, customer correspondence and contracts contain personal data. Ask whether each document group is truly needed; not connecting an unnecessary group is the strongest safeguard.
  • Processors: When documents with personal data reach a model provider, that provider acts as a processor and needs a data processing agreement under Article 28 of the GDPR.
  • International transfers: If data leaves the EU or EEA, Chapter V of the GDPR requires an adequacy decision or appropriate safeguards such as standard contractual clauses. UK and US based teams should check their own transfer and privacy rules with counsel.
  • Sensitive groups: Health records and similar files are either kept out of the index or processed only by a local model with narrow access.

Article 4 of the EU AI Act, as amended in 2026, asks organizations that deploy AI systems to take measures to support their staff's AI literacy. We treat this less as a compliance box and more as a precondition for using the assistant well; corporate AI training can close that gap for your teams.

This is not legal advice; finalize the transfer basis and employee notice with your legal adviser.

Run the pilot with written criteria

A pilot runs with one team, a limited document set and success criteria written down beforehand. Without criteria, nobody can make a decision when the pilot ends.

  1. Write the scope: Which team, which document groups and which question types, plus a list of what is out of scope.
  2. Collect the question set: Gather real questions and correct answers with their source passage from the team, mixing easy, hard and not in the documents cases.
  3. Measure the baseline: Time sample lookups today; a working hours calculator turns that into weekly team hours.
  4. Finish internal testing: Run the question set and test citations, the permission filter and the deleted document scenario separately.
  5. Open to part of the team: A subset starts using it, and feedback plus unanswered questions are reviewed weekly.
  6. Hold the decision meeting: If criteria are met, add a second team or document group; if not, use the logs to locate the problem in documents, retrieval or the model.

Tie pilot length to question volume rather than the calendar; a pilot few people used is simply inconclusive. When you expand, each new document group arrives with its own questions, so older groups are retested every time.

Five indicators that show real value

Value is tracked with indicators that separately measure whether retrieval finds the right passage, whether answers stay faithful to sources and whether the team actually adopts the tool, not with a single satisfaction score.

  • Retrieval hit rate: For each question in the set, is the correct passage among the top results? If this is low, the problem lies in chunking and search, not the model.
  • Faithfulness: Is every claim in the answer actually written in the cited passage? Measured by human review on a sample.
  • Correct "I don't know": For questions with no answer in the documents, how often does the assistant say so instead of inventing one?
  • Return usage: Do pilot users come back after the first week? This separates curiosity from habit.
  • Questions to experts: Are routine questions to senior staff declining, tracked with a short survey or by counting messages in an internal support channel?

Recalculate these with the same question set whenever the index or model changes, since one measurement cannot reveal slow decay. If you want a financial figure from RAG development, base it on your own measured search time and headcount, not on ratios borrowed from other companies.

Where retrieval augmented generation falls short

RAG is strong at finding and relaying what documents say, and weak at inferring what they do not say, counting across an archive and resolving contradictions. Knowing this upfront sets the right expectations.

  • Aggregate questions: "How many contracts this year include a penalty clause" requires scanning the whole archive; RAG retrieves only a handful of relevant passages and may produce an incomplete number. Reporting tools handle such needs.
  • Conflicting documents: When two valid documents disagree, the assistant should surface the conflict rather than pick one; the lasting fix belongs to the document owner.
  • Fluent but wrong: A model can build a convincing sentence on an incomplete passage. Citations reduce this risk without removing it, so users should open the source on critical topics.
  • Hidden instructions: Text planted in an external file may try to steer the model; read access only and no outbound actions are essential for that reason.
  • Stale permissions: If someone changed roles and their old access was never removed at the source, the assistant keeps honoring it.

Overreliance is the quietest risk: the more often the assistant is right, the less people check sources. Handover training therefore shows which question types it handles poorly.

Common mistakes and what to do instead

Most RAG development mistakes are about scope and order, not technology. Each of these six can be prevented before setup.

  • Connecting the whole drive at once: Drafts, copies and old versions enter the index too. Start with one document group and add others only after cleaning them.
  • Judging by demo questions: A vendor's own questions always look good. Insist on a set built from your team's real questions, including ones the documents cannot answer.
  • Adding permissions later: Retrofitting a permission filter may force an index rebuild. Make access control part of the index from day one.
  • No document owners: If nobody is responsible for updates, the assistant serves outdated facts within months. Put a name on every source group.
  • Ignoring shadow AI use: Staff keep pasting company files into personal chat accounts. When you launch, publish a short usage rule explaining why that should stop.
  • Measuring only satisfaction: Thumbs up ratings hide wrong answers. Track question set results, faithfulness and the unanswered list together.

What these share is treating the assistant as a one time install rather than a system that changes with your documents.

Questions to ask a RAG partner

When comparing RAG development proposals, look past the demo to how the team catches errors, carries permissions and what it leaves you at handover. Ask for written answers to these:

  • Permissions: How and how quickly does a permission change at the source reach the index, and is a deleted document test part of acceptance?
  • Testing: Whose questions form the question set, are retrieval and answers measured separately, and will you see results after every update?
  • Data path: Which passages go to which provider and region, and can they show documentation that the provider does not train on API data?
  • Ownership: In whose name are accounts, index settings, prompts and the question set; could another team take the system over?
  • Maintenance: How are new document types, model changes and quality drops handled after launch?

Asking for references is fair. We have not yet delivered a knowledge assistant for a client we can name; our automation, software and web work is on our references page. So we present our method in writing with pilot criteria and suggest you judge it with your own questions. Build, connectors and upkeep run within our AI automation services; a proposal that leaves maintenance vague should weigh heavily in your comparison.

Making the decision and the first step

If you answered yes to at least three readiness questions and one team clearly carries a load of repeated questions, you are in a good position for a pilot; if not, document cleanup comes first. To make the first conversation productive, gather:

  • The team losing the most time and three to five typical questions from it
  • The systems and document types where those answers live
  • Document groups you treat as sensitive and whether they may leave your infrastructure
  • Where people will use the assistant: browser, Teams, Slack or an internal portal
  • The criterion that would make the pilot a success for you

With this, a short call is enough to agree on pilot scope, the data path and high risk document groups for your RAG development project. You can request that call through our contact page.

Cost depends on document volume, number of source systems, permission complexity and hosting choice. Fixed scope options for discovery and the first workflow are in the AI automation pricing section, and a written quote follows once scope is clear. If the readiness test says not yet, we tell you plainly, and the outcome of the call becomes a cleanup list for your documentation.