DetermindsStart a project
Services
Industries
Case studies
Products
Company
Start a project

Custom RAG system case study: AI search across 10,000+ PDFs

Determinds built a custom RAG system for a policy research organization working with more than 10,000 United Nations General Assembly resolutions, public documents such as A/RES/76/300. The goal was clear from the start: “I want my team to ask a question and get a cited answer in seconds, instead of spending an afternoon hunting through PDFs.”

Client
A policy research organization
Sector
NGO
Offering
Custom Software Solutions
Documents
10,000+ United Nations General Assembly resolutions

In short

Updated

Determinds built a custom RAG system that gives a policy research organization AI-powered search across 10,000+ UN General Assembly documents, cutting typical research lookups from 30 to 90 minutes to under 10 seconds, with cited answers.

Documents
10,000+ UN General Assembly resolutions, as PDFs
Options evaluated
Keyword search, fine-tuning a custom LLM, and Retrieval-Augmented Generation (RAG). RAG was chosen for its cost, accuracy, and ease of maintenance
Research lookups
From 30 to 90 minutes of manual search to under 10 seconds
Answers
Cited, grounded in the source documents
Also works for
Legal archives, internal knowledge bases, technical manuals, financial filings, and customer support repositories
Diagram of a custom RAG system pipeline: PDF documents flow through ingest, chunk, embed, and vector database stages to produce a cited AI answer.

Determinds designed and built a custom RAG system (Retrieval-Augmented Generation) that delivers exactly that. The same approach is now the most cost-effective way for any organization with a large document collection to turn it into a conversational, AI-powered knowledge base, whether that collection is a document archive, a contract library, a regulatory corpus, or a customer support history.

01  The challenge

The challenge of searching 10,000+ UN resolutions

Our client is a policy research organization whose analysts cite UN General Assembly resolutions constantly, in briefings, position papers, legal opinions, and historical research. The archive grows every year, and each new session adds hundreds of formally numbered documents in multiple languages.

Before the new system, research followed four steps:

  • An analyst received a research question, such as “What has the General Assembly said about the right to a clean and healthy environment since 2018?”).
  • The analyst opened the UN documents portal, ran keyword searches, and downloaded a stack of PDFs.
  • They skimmed each PDF, copied relevant paragraphs into a notes document, and kept track of which resolution said what.
  • A 30-minute question routinely became a half-day project.

Keyword search on the source portal worked when the analyst knew the exact phrase a document used. Policy language, however, is dense and varied. Resolutions use terms such as “Member States are encouraged to” or “calls upon,” and the same idea can be phrased a dozen different ways across documents. A keyword query for “climate financing” would miss a resolution that referred to “financial support for mitigation in developing countries”, even though that was exactly the relevant text.

The client wanted analysts to ask a question in natural language and receive a grounded, cited answer, drawn directly from the resolutions themselves and always traceable to its source. In policy work, a verifiable source is essential to every answer, so this was the core requirement.

02  Options evaluated

Three approaches to AI document search

Before writing any code, we walked the client through the three realistic approaches to building an AI document search system in 2026, and the trade-offs of each.

Option 1

Better keyword search (Elasticsearch / Algolia)

A modern search engine such as Elasticsearch or Algolia is fast, mature, and inexpensive. It would give the client filters, facets, and fuzzy matching, a clear upgrade over the UN portal’s built-in search. It still matches words rather than understanding the question, however. It looks for tokens that appear in the text, so instead of answering “What position has been taken on X?” it returns a list of documents for the analyst to read.

It makes the archive easier to navigate, and analysts still read each document.

Option 2

Fine-tuning a custom language model

Fine-tuning takes a base AI model, such as a Llama or Mistral variant, and continues its training on the client’s documents until the model absorbs them. The idea is appealing, because the AI would know the corpus in depth. For a use case like this one, fine-tuning has three serious drawbacks:

  • Cost. Fine-tuning a base model on a corpus this size typically reaches the high four- to five-figure range in compute alone, based on published GPU pricing from providers such as AWS SageMaker and Lambda Labs, and the cost repeats every time the archive updates.
  • Hallucinations. A fine-tuned model still generates text from patterns, and it can confidently invent a citation that does not exist. For a research organization, that rules it out.
  • No verifiable sources. Even when the model is right, it cannot show which resolution the answer came from, so the analyst still has to verify everything.
Option 3

A custom RAG system

Retrieval-Augmented Generation (RAG) is a hybrid approach. Instead of training the documents into the AI model itself, we keep them in a separate, searchable index. When a user asks a question, the system first retrieves the most relevant passages from the archive, then passes them to a language model with instructions to write a grounded answer based only on what was retrieved.

The result works like a conversation with a skilled research assistant who has every document open on the desk and cites a real source for every answer.

According to the original 2020 Meta AI research paper that introduced the technique, RAG systems outperform both pure search and pure generation on knowledge-intensive tasks, and modern embedding models have widened that gap further.

03  Why RAG

Why we recommended a custom RAG system

For this client, and for most organizations we speak with that want AI on top of their documents, RAG was the clear choice. This is the side-by-side comparison we presented:

Comparing AI document search approaches

ApproachSetup costOngoing costCitationsEasy to updateRisk of hallucination
Keyword searchLowLowYes (links only)YesNone, and no direct answers
Fine-tuned LLMVery highHigh (retraining)NoNo (full retrain)High
Custom RAG systemModerateLow to moderateYes (passage-level)Yes (add documents anytime)Low (grounded answers)

Three factors decided it for our client:

Every answer comes with citations

The system returns the exact resolution, paragraph, and page behind each answer, so an analyst can verify a claim in one click.

Adding documents is simple

When the next UN session publishes 400 new resolutions, we ingest them in a single overnight batch, with no retraining required.

The economics work at any scale

The cost per question comes mainly from one retrieval and one language-model call, well under a cent at current OpenAI and Anthropic API rates. The system pays for itself in the first week an analyst no longer spends half a day on a search.

04  How we built it

How we built the custom RAG system, step by step

The custom RAG pipeline Determinds built for the UN resolutions archive, from ingestion to cited answer.

A custom RAG system has four core components, and each one can be explained in plain terms without an engineering background. Together they form the pipeline at a glance:

Ingestion: turning PDFs into clean text

UN resolutions are published as PDFs, often scanned and sometimes with footnotes, headers, and multi-column layouts, so the first requirement was clean text. We built an ingestion pipeline that pulls each PDF, runs scanned pages through an OCR layer, strips boilerplate such as page numbers and repeated headers, and tags each document with metadata: resolution number, session, date, language, and topic. This step sets the quality of every answer the system gives, because accurate answers depend on clean source text.

PDFs → OCR → clean text + metadata

Chunking: dividing documents into focused passages

A 40-page document is too long to hand to an AI model to find one relevant sentence, so we divide each document into smaller, self-contained chunks, typically 300 to 800 words each, that the system can match to a question. For UN resolutions, we chunked at the paragraph level, since resolutions are conveniently numbered by paragraph, with a small overlap between chunks so that an idea split across two paragraphs stays intact. Each chunk carries the metadata of its parent document, so every retrieved chunk can be traced to exactly where it came from.

Split into passages of 300 to 800 words

Embeddings: capturing what each chunk means

For each chunk, an embedding model, in this project OpenAI’s text-embedding-3-large, converts the text into a list of numbers (a 3,072-dimension vector) that acts as a mathematical fingerprint of its meaning. Two paragraphs that express the same idea in different words have fingerprints close to each other in mathematical space, even when they share no keywords. That is how the system recognizes that “climate financing” and “financial support for mitigation” describe the same concept, which is the foundation of semantic search.

Each chunk converted to a meaning vector

Vector database: where the fingerprints are stored

We store every chunk’s fingerprint in a vector database, a specialized search index built to find the closest matches across millions of numbers in milliseconds. For this project we used a managed vector store, which keeps running costs low, and the system is specified for 100,000 documents.

Indexed for millisecond retrieval

User question → Embed → Retrieve top chunks → LLM writes cited answer

When an analyst asks a question, the same embedding model converts it into a fingerprint, the vector database returns the 5 to 10 most relevant chunks, and a large language model, in our build OpenAI’s GPT-5.5, writes a clear, cited answer grounded strictly in those retrieved chunks. End to end, the analyst receives an answer in about 2 to 4 seconds..

The stack at a glance: text-embedding-3-large for embeddings, a managed vector store for retrieval, and GPT-5.5 for grounded answer generation, all connected through a custom ingestion pipeline and a chat interface hosted in the client’s own environment.

The system also includes guardrails that decline questions the documents do not answer, a feedback loop that lets analysts flag weak responses, language detection for multilingual resolutions, and an admin panel for ingesting new documents. These four steps form the core of every custom RAG system we build.

05  Video demo

Video demo of the RAG system

The 2-minute walkthrough shows the system handling a real research question, from typing a natural-language query to receiving a grounded answer in seconds and clicking through to the exact UN resolution and paragraph it cited.

06  Business results

Business results in research time, cost, and confidence

Every custom software project is measured by its business outcome. After launch, the custom RAG system delivered these results for the client:

  • Research time per question fell from 30 to 90 minutes to under 10 seconds. A typical analyst now completes in a morning what used to take a full day.
  • Citation confidence increased. Every answer links to the exact resolution and paragraph, and the team trusts the output enough to use it directly in briefings.
  • New analysts onboard faster. Junior team members no longer need a year of institutional memory to navigate the archive, because the system surfaces the right context on demand.
  • The archive grows without extra effort. Each new UN session is ingested overnight, with no retraining, manual indexing, or outside consultants required.
  • Running costs stay minimal. The total monthly cost, covering the vector database, the embedding API, and language model calls, is lower than the team previously spent on a single contractor day.

Across the RAG projects we have delivered, the value is consistent. AI takes on the most repetitive part of an analyst’s work, so analysts can focus on the strategic, judgment-heavy work their role exists for.

07  Where it fits

Where a custom RAG system can work for your business

The UN resolutions project is one specific use case, and the underlying pattern applies widely. If your organization has any of the following, a custom RAG system is very likely worth a 30-minute conversation:

A legal or contract library

Clauses, precedents, NDAs, and vendor agreements that lawyers re-read constantly.

An internal knowledge base

Confluence pages, SOPs, training materials, and onboarding documents your team needs to find quickly.

Regulatory or compliance documents

FDA filings, financial regulations, audit trails, and ISO certifications.

Customer support history

Tickets, transcripts, and resolved cases that a support agent can search in seconds instead of minutes.

Technical manuals and product documentation

Engineering specs, repair manuals, and integration documents across multiple product lines.

Financial reports and analyst archives

Investor decks, earnings transcripts, and market research.

The economics improve as the document count grows. Below a few hundred documents, a well-organized folder and a good search bar are usually enough. Above a few thousand, the collection becomes a knowledge management challenge that benefits from a dedicated system, and a custom RAG system typically pays for itself within the first quarter.

08  8-point self-check

Eight signs of a RAG-ready archive

Each statement is a sign that an archive is ready for RAG. Where five or more apply, a custom RAG system will very likely pay back its build cost within the first quarter.

  • We have more than 1,000 documents (PDFs, Word files, internal pages, transcripts) that staff search regularly.
  • A typical research or lookup question takes a person more than 15 minutes today.
  • The same questions are asked repeatedly by different team members.
  • Keyword search misses results when the phrasing differs from the source text.
  • Answers must include a verifiable citation (legal, regulatory, research, and compliance contexts).
  • The document set grows continuously (new contracts, filings, support tickets, and releases).
  • Subject-matter experts spend hours on lookups that a junior colleague could handle with the right tool.
  • Documents must stay inside your environment (private cloud, on-premises, or a regulated industry).

Determinds builds these systems end to end, from the ingestion pipeline to the chat interface, and deploys them on infrastructure you own, so your documents stay inside your environment. The fastest way to confirm the fit is a conversation with our team. We will review your document set and use case, and recommend whether RAG is the right tool or whether a simpler system will serve you better.

09  Conclusion

From document archive to conversational knowledge base

For many organizations, the document archive built up over years of work is one of their most valuable and least used assets, mainly because nobody has time to read it. A custom RAG system turns that archive into a conversation: ask a question, receive a cited answer, and move on with your day.

Our policy research client went from needing an afternoon to needing ten seconds, and the same shift is available to any organization ready to invest a few weeks in the right architecture. Our team would welcome that conversation with yours.

10  Frequently asked

Frequently asked questions

What is a custom RAG system in simple terms?

A custom RAG (Retrieval-Augmented Generation) system is an AI tool that answers natural-language questions about your own documents with cited, grounded answers. It first retrieves the most relevant passages from your archive, then uses a language model to write an answer based only on what it retrieved, so you get the conversational power of AI with every answer anchored in your sources.

How is RAG different from just using ChatGPT?

A custom RAG system answers from your own private documents, while ChatGPT and similar general-purpose models only know what was in their training data. The RAG system connects an AI model directly to your files, so every answer is drawn from your contracts, policies, manuals, or research and comes with a citation pointing back to the source.

How long does it take to build a custom RAG system?

A production-ready custom RAG system typically takes 4 to 10 weeks, depending on the document set, the complexity of the source files (clean text, scanned PDFs, or mixed media), and the user interface required. A working prototype for internal demos is often achievable in 2 to 3 weeks.

How much does a custom RAG system cost to run?

Running costs depend on usage volume. For most mid-sized document archives (10,000 to 100,000 documents) with moderate query volumes, monthly costs range from a few hundred to a few thousand dollars, covering vector database hosting, embedding APIs, and language model calls. The cost per question is typically a fraction of a cent.

Will my documents be safe if I use a custom RAG system?

Yes, when the system is designed for it. Determinds builds custom RAG systems on infrastructure you own, with your documents encrypted at rest and never used to train external AI models. For clients in regulated industries, we deploy the entire stack inside the client’s own cloud account, so every document and query stays within their environment.

Can a RAG system handle documents in multiple languages?

Yes. Modern multilingual embedding models, including OpenAI’s text-embedding-3-large (the model used in this project), cover roughly 100 languages and can match a question in one language to relevant passages in another. The UN resolutions archive, for example, includes documents in English, French, and Arabic, and analysts can ask questions in any of those languages and receive answers grounded in the appropriate source text.

Sources

References

  1. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”Meta AI / arXiv, 2020. The foundational RAG paper.
  2. United Nations, General Assembly Resolution A/RES/76/300Example source document.
  3. AWS SageMaker pricingBasis for fine-tuning cost estimates.
  4. Lambda Labs GPU cloud pricingBasis for fine-tuning cost estimates.
  5. OpenAI API pricingBasis for per-query cost estimates.
  6. Anthropic API pricingBasis for per-query cost estimates.
  7. OpenAI Embeddings documentationReference for the text-embedding-3-large model used in this project.

Free discovery call

Turn your document archive into searchable, cited answers

Book a free 30-minute discovery call with a Determinds engineer. We will review your use case and set out what a custom RAG system would cost, how long it would take, and whether it is the right fit.