Custom RAG system case study: AI search across 10,000+ PDFs
Determinds built a custom RAG system for a policy research organization working with more than 10,000 United Nations General Assembly resolutions, public documents such as A/RES/76/300. The goal was clear from the start: “I want my team to ask a question and get a cited answer in seconds, instead of spending an afternoon hunting through PDFs.”
- Client
- A policy research organization
- Sector
- NGO
- Offering
- Custom Software Solutions
- Documents
- 10,000+ United Nations General Assembly resolutions
In short
Updated
Determinds built a custom RAG system that gives a policy research organization AI-powered search across 10,000+ UN General Assembly documents, cutting typical research lookups from 30 to 90 minutes to under 10 seconds, with cited answers.
- Documents
- 10,000+ UN General Assembly resolutions, as PDFs
- Options evaluated
- Keyword search, fine-tuning a custom LLM, and Retrieval-Augmented Generation (RAG). RAG was chosen for its cost, accuracy, and ease of maintenance
- Research lookups
- From 30 to 90 minutes of manual search to under 10 seconds
- Answers
- Cited, grounded in the source documents
- Also works for
- Legal archives, internal knowledge bases, technical manuals, financial filings, and customer support repositories

Determinds designed and built a custom RAG system (Retrieval-Augmented Generation) that delivers exactly that. The same approach is now the most cost-effective way for any organization with a large document collection to turn it into a conversational, AI-powered knowledge base, whether that collection is a document archive, a contract library, a regulatory corpus, or a customer support history.
Contents
Table of contents
The challenge of searching 10,000+ UN resolutions
Three approaches to AI document search
Why we recommended a custom RAG system
How we built the custom RAG system, step by step
Video demo of the RAG system
Business results in research time, cost, and confidence
Where a custom RAG system can work for your business
Frequently asked questions
01 The challenge
The challenge of searching 10,000+ UN resolutions
Our client is a policy research organization whose analysts cite UN General Assembly resolutions constantly, in briefings, position papers, legal opinions, and historical research. The archive grows every year, and each new session adds hundreds of formally numbered documents in multiple languages.
Before the new system, research followed four steps:
- An analyst received a research question, such as “What has the General Assembly said about the right to a clean and healthy environment since 2018?”).
- The analyst opened the UN documents portal, ran keyword searches, and downloaded a stack of PDFs.
- They skimmed each PDF, copied relevant paragraphs into a notes document, and kept track of which resolution said what.
- A 30-minute question routinely became a half-day project.
Keyword search on the source portal worked when the analyst knew the exact phrase a document used. Policy language, however, is dense and varied. Resolutions use terms such as “Member States are encouraged to” or “calls upon,” and the same idea can be phrased a dozen different ways across documents. A keyword query for “climate financing” would miss a resolution that referred to “financial support for mitigation in developing countries”, even though that was exactly the relevant text.
The client wanted analysts to ask a question in natural language and receive a grounded, cited answer, drawn directly from the resolutions themselves and always traceable to its source. In policy work, a verifiable source is essential to every answer, so this was the core requirement.
02 Options evaluated
Three approaches to AI document search
Before writing any code, we walked the client through the three realistic approaches to building an AI document search system in 2026, and the trade-offs of each.
Better keyword search (Elasticsearch / Algolia)
A modern search engine such as Elasticsearch or Algolia is fast, mature, and inexpensive. It would give the client filters, facets, and fuzzy matching, a clear upgrade over the UN portal’s built-in search. It still matches words rather than understanding the question, however. It looks for tokens that appear in the text, so instead of answering “What position has been taken on X?” it returns a list of documents for the analyst to read.
It makes the archive easier to navigate, and analysts still read each document.
Fine-tuning a custom language model
Fine-tuning takes a base AI model, such as a Llama or Mistral variant, and continues its training on the client’s documents until the model absorbs them. The idea is appealing, because the AI would know the corpus in depth. For a use case like this one, fine-tuning has three serious drawbacks:
- Cost. Fine-tuning a base model on a corpus this size typically reaches the high four- to five-figure range in compute alone, based on published GPU pricing from providers such as AWS SageMaker and Lambda Labs, and the cost repeats every time the archive updates.
- Hallucinations. A fine-tuned model still generates text from patterns, and it can confidently invent a citation that does not exist. For a research organization, that rules it out.
- No verifiable sources. Even when the model is right, it cannot show which resolution the answer came from, so the analyst still has to verify everything.
A custom RAG system
Retrieval-Augmented Generation (RAG) is a hybrid approach. Instead of training the documents into the AI model itself, we keep them in a separate, searchable index. When a user asks a question, the system first retrieves the most relevant passages from the archive, then passes them to a language model with instructions to write a grounded answer based only on what was retrieved.
The result works like a conversation with a skilled research assistant who has every document open on the desk and cites a real source for every answer.
According to the original 2020 Meta AI research paper that introduced the technique, RAG systems outperform both pure search and pure generation on knowledge-intensive tasks, and modern embedding models have widened that gap further.
03 Why RAG
Why we recommended a custom RAG system
For this client, and for most organizations we speak with that want AI on top of their documents, RAG was the clear choice. This is the side-by-side comparison we presented:
Comparing AI document search approaches
| Approach | Setup cost | Ongoing cost | Citations | Easy to update | Risk of hallucination |
|---|---|---|---|---|---|
| Keyword search | Low | Low | Yes (links only) | Yes | None, and no direct answers |
| Fine-tuned LLM | Very high | High (retraining) | No | No (full retrain) | High |
| Custom RAG system | Moderate | Low to moderate | Yes (passage-level) | Yes (add documents anytime) | Low (grounded answers) |
Three factors decided it for our client:
Every answer comes with citations
The system returns the exact resolution, paragraph, and page behind each answer, so an analyst can verify a claim in one click.
Adding documents is simple
When the next UN session publishes 400 new resolutions, we ingest them in a single overnight batch, with no retraining required.
The economics work at any scale
The cost per question comes mainly from one retrieval and one language-model call, well under a cent at current OpenAI and Anthropic API rates. The system pays for itself in the first week an analyst no longer spends half a day on a search.
04 How we built it
How we built the custom RAG system, step by step
The custom RAG pipeline Determinds built for the UN resolutions archive, from ingestion to cited answer.
A custom RAG system has four core components, and each one can be explained in plain terms without an engineering background. Together they form the pipeline at a glance:
Ingestion: turning PDFs into clean text
UN resolutions are published as PDFs, often scanned and sometimes with footnotes, headers, and multi-column layouts, so the first requirement was clean text. We built an ingestion pipeline that pulls each PDF, runs scanned pages through an OCR layer, strips boilerplate such as page numbers and repeated headers, and tags each document with metadata: resolution number, session, date, language, and topic. This step sets the quality of every answer the system gives, because accurate answers depend on clean source text.
Chunking: dividing documents into focused passages
A 40-page document is too long to hand to an AI model to find one relevant sentence, so we divide each document into smaller, self-contained chunks, typically 300 to 800 words each, that the system can match to a question. For UN resolutions, we chunked at the paragraph level, since resolutions are conveniently numbered by paragraph, with a small overlap between chunks so that an idea split across two paragraphs stays intact. Each chunk carries the metadata of its parent document, so every retrieved chunk can be traced to exactly where it came from.
Embeddings: capturing what each chunk means
For each chunk, an embedding model, in this project OpenAI’s text-embedding-3-large, converts the text into a list of numbers (a 3,072-dimension vector) that acts as a mathematical fingerprint of its meaning. Two paragraphs that express the same idea in different words have fingerprints close to each other in mathematical space, even when they share no keywords. That is how the system recognizes that “climate financing” and “financial support for mitigation” describe the same concept, which is the foundation of semantic search.
Vector database: where the fingerprints are stored
We store every chunk’s fingerprint in a vector database, a specialized search index built to find the closest matches across millions of numbers in milliseconds. For this project we used a managed vector store, which keeps running costs low, and the system is specified for 100,000 documents.
User question → Embed → Retrieve top chunks → LLM writes cited answer
When an analyst asks a question, the same embedding model converts it into a fingerprint, the vector database returns the 5 to 10 most relevant chunks, and a large language model, in our build OpenAI’s GPT-5.5, writes a clear, cited answer grounded strictly in those retrieved chunks. End to end, the analyst receives an answer in about 2 to 4 seconds..
The stack at a glance: text-embedding-3-large for embeddings, a managed vector store for retrieval, and GPT-5.5 for grounded answer generation, all connected through a custom ingestion pipeline and a chat interface hosted in the client’s own environment.
The system also includes guardrails that decline questions the documents do not answer, a feedback loop that lets analysts flag weak responses, language detection for multilingual resolutions, and an admin panel for ingesting new documents. These four steps form the core of every custom RAG system we build.
05 Video demo
Video demo of the RAG system
The 2-minute walkthrough shows the system handling a real research question, from typing a natural-language query to receiving a grounded answer in seconds and clicking through to the exact UN resolution and paragraph it cited.
06 Business results
Business results in research time, cost, and confidence
Every custom software project is measured by its business outcome. After launch, the custom RAG system delivered these results for the client:
- Research time per question fell from 30 to 90 minutes to under 10 seconds. A typical analyst now completes in a morning what used to take a full day.
- Citation confidence increased. Every answer links to the exact resolution and paragraph, and the team trusts the output enough to use it directly in briefings.
- New analysts onboard faster. Junior team members no longer need a year of institutional memory to navigate the archive, because the system surfaces the right context on demand.
- The archive grows without extra effort. Each new UN session is ingested overnight, with no retraining, manual indexing, or outside consultants required.
- Running costs stay minimal. The total monthly cost, covering the vector database, the embedding API, and language model calls, is lower than the team previously spent on a single contractor day.
Across the RAG projects we have delivered, the value is consistent. AI takes on the most repetitive part of an analyst’s work, so analysts can focus on the strategic, judgment-heavy work their role exists for.
07 Where it fits
Where a custom RAG system can work for your business
The UN resolutions project is one specific use case, and the underlying pattern applies widely. If your organization has any of the following, a custom RAG system is very likely worth a 30-minute conversation:
A legal or contract library
Clauses, precedents, NDAs, and vendor agreements that lawyers re-read constantly.
An internal knowledge base
Confluence pages, SOPs, training materials, and onboarding documents your team needs to find quickly.
Regulatory or compliance documents
FDA filings, financial regulations, audit trails, and ISO certifications.
Customer support history
Tickets, transcripts, and resolved cases that a support agent can search in seconds instead of minutes.
Technical manuals and product documentation
Engineering specs, repair manuals, and integration documents across multiple product lines.
Financial reports and analyst archives
Investor decks, earnings transcripts, and market research.
The economics improve as the document count grows. Below a few hundred documents, a well-organized folder and a good search bar are usually enough. Above a few thousand, the collection becomes a knowledge management challenge that benefits from a dedicated system, and a custom RAG system typically pays for itself within the first quarter.
08 8-point self-check
Eight signs of a RAG-ready archive
Each statement is a sign that an archive is ready for RAG. Where five or more apply, a custom RAG system will very likely pay back its build cost within the first quarter.
- We have more than 1,000 documents (PDFs, Word files, internal pages, transcripts) that staff search regularly.
- A typical research or lookup question takes a person more than 15 minutes today.
- The same questions are asked repeatedly by different team members.
- Keyword search misses results when the phrasing differs from the source text.
- Answers must include a verifiable citation (legal, regulatory, research, and compliance contexts).
- The document set grows continuously (new contracts, filings, support tickets, and releases).
- Subject-matter experts spend hours on lookups that a junior colleague could handle with the right tool.
- Documents must stay inside your environment (private cloud, on-premises, or a regulated industry).
Determinds builds these systems end to end, from the ingestion pipeline to the chat interface, and deploys them on infrastructure you own, so your documents stay inside your environment. The fastest way to confirm the fit is a conversation with our team. We will review your document set and use case, and recommend whether RAG is the right tool or whether a simpler system will serve you better.
09 Conclusion
From document archive to conversational knowledge base
For many organizations, the document archive built up over years of work is one of their most valuable and least used assets, mainly because nobody has time to read it. A custom RAG system turns that archive into a conversation: ask a question, receive a cited answer, and move on with your day.
Our policy research client went from needing an afternoon to needing ten seconds, and the same shift is available to any organization ready to invest a few weeks in the right architecture. Our team would welcome that conversation with yours.
10 Frequently asked
Frequently asked questions
What is a custom RAG system in simple terms?
A custom RAG (Retrieval-Augmented Generation) system is an AI tool that answers natural-language questions about your own documents with cited, grounded answers. It first retrieves the most relevant passages from your archive, then uses a language model to write an answer based only on what it retrieved, so you get the conversational power of AI with every answer anchored in your sources.
How is RAG different from just using ChatGPT?
A custom RAG system answers from your own private documents, while ChatGPT and similar general-purpose models only know what was in their training data. The RAG system connects an AI model directly to your files, so every answer is drawn from your contracts, policies, manuals, or research and comes with a citation pointing back to the source.
How long does it take to build a custom RAG system?
A production-ready custom RAG system typically takes 4 to 10 weeks, depending on the document set, the complexity of the source files (clean text, scanned PDFs, or mixed media), and the user interface required. A working prototype for internal demos is often achievable in 2 to 3 weeks.
How much does a custom RAG system cost to run?
Running costs depend on usage volume. For most mid-sized document archives (10,000 to 100,000 documents) with moderate query volumes, monthly costs range from a few hundred to a few thousand dollars, covering vector database hosting, embedding APIs, and language model calls. The cost per question is typically a fraction of a cent.
Will my documents be safe if I use a custom RAG system?
Yes, when the system is designed for it. Determinds builds custom RAG systems on infrastructure you own, with your documents encrypted at rest and never used to train external AI models. For clients in regulated industries, we deploy the entire stack inside the client’s own cloud account, so every document and query stays within their environment.
Can a RAG system handle documents in multiple languages?
Yes. Modern multilingual embedding models, including OpenAI’s text-embedding-3-large (the model used in this project), cover roughly 100 languages and can match a question in one language to relevant passages in another. The UN resolutions archive, for example, includes documents in English, French, and Arabic, and analysts can ask questions in any of those languages and receive answers grounded in the appropriate source text.
Sources
References
- Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”Meta AI / arXiv, 2020. The foundational RAG paper.
- United Nations, General Assembly Resolution A/RES/76/300Example source document.
- AWS SageMaker pricingBasis for fine-tuning cost estimates.
- Lambda Labs GPU cloud pricingBasis for fine-tuning cost estimates.
- OpenAI API pricingBasis for per-query cost estimates.
- Anthropic API pricingBasis for per-query cost estimates.
- OpenAI Embeddings documentationReference for the text-embedding-3-large model used in this project.
Free discovery call
Turn your document archive into searchable, cited answers
Book a free 30-minute discovery call with a Determinds engineer. We will review your use case and set out what a custom RAG system would cost, how long it would take, and whether it is the right fit.