Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,6 +76,25 @@ python examples/rag_demo.py

This repository is offline-first and intentionally uses deterministic mock embeddings rather than paid external embedding providers. The examples and tests are designed to run entirely on local fixtures. No secret, credential, or hosted service is required for standard execution.

## Workshop Corpus

The fixture corpus is intentionally built as a compact but realistic RAG workshop. Instead of three tiny text files, the project now includes six original Markdown documents that cover the main concepts of retrieval:

- a project overview and RAG fundamentals
- document loading and indexing
- chunking and segmentation
- embeddings and vector search
- MCP tool contracts and structured tool discovery
- evaluation and end-to-end retrieval workflows

These documents are designed to support questions such as:

- "What is retrieval augmented generation?"
- "How does a retriever choose the best chunks?"
- "What does an MCP tool do?"

The loader supports both `.txt` and `.md` files and preserves deterministic ordering, while the workshop corpus emphasizes the Markdown workflow that is common in documentation-heavy RAG use cases.

## Architecture Explanation

### Documents
Expand Down
21 changes: 14 additions & 7 deletions examples/rag_demo.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,13 +33,20 @@ def main() -> None:
chunks=chunks,
)

query = "MCP tools and retrieval basics"
results = retriever.retrieve(query, top_k=3)
print(f"Query: {query}")
print("Results:")
for index, result in enumerate(results, start=1):
print(f" {index}. {result.source} (score={result.score:.4f})")
print(f" {result.text[:120]}...")
queries = [
"What is retrieval augmented generation?",
"What does an MCP tool do?",
"How does a vector store rank relevant chunks?",
]

for query in queries:
results = retriever.retrieve(query, top_k=3)
print(f"Query: {query}")
print("Results:")
for index, result in enumerate(results, start=1):
print(f" {index}. {result.source} (score={result.score:.4f})")
print(f" {result.text[:120]}...")
print()


if __name__ == "__main__":
Expand Down
9 changes: 9 additions & 0 deletions fixtures/documents/01-overview.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Retrieval-Augmented Generation Overview

Retrieval-augmented generation, or RAG, combines a searchable knowledge base with a language model. Instead of asking the model to memorize everything in its weights, the system first retrieves the most relevant evidence from a corpus and then uses that evidence as context.

A typical RAG pipeline starts with a document loader, which reads source files from disk or a remote store. The loader converts each source into a normalized document object with metadata such as filename, language, and source path. The next stage splits each document into chunks so that retrieval can focus on smaller, semantically meaningful passages.

Once chunks are created, an embedding model maps them into a vector space. A FAISS index stores those vectors for efficient similarity search. At query time, the system embeds the user question, searches the index, and returns the most relevant chunks. The selected passages are then supplied to a downstream model for final reasoning or synthesis.

This project is intentionally small, transparent, and offline-first. It teaches the mechanics behind retrieval without hiding them behind a large framework.
9 changes: 9 additions & 0 deletions fixtures/documents/02-loading-and-indexing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Document Loading and Indexing

Document loading is the first step in any RAG system. The loader reads text files, normalizes whitespace, and preserves metadata that will later be used for debugging and traceability. A realistic loader should be deterministic, support UTF-8 content, and ignore files that are not appropriate for retrieval.

In this workshop, the loader accepts plain text and Markdown files and produces a `Document` object for each file. A normalized document record includes the document ID, the raw text, and metadata such as the filename and source path. The loader sorts files consistently so that the corpus order remains stable across repeated runs.

After loading, the indexer prepares those documents for search. Each chunk is turned into an embedding, and those embeddings are inserted into a FAISS index. A vector store preserves the relationship between a chunk and its metadata, which allows the retriever to return the matching text and its source document.

This separation of concerns is important. Loading extracts content, indexing prepares the data for search, and retrieval answers the user question by comparing the query embedding to the stored vectors.
9 changes: 9 additions & 0 deletions fixtures/documents/03-chunking.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Chunking Strategies

Chunking is the process of splitting a document into smaller segments before indexing. A chunk should be large enough to preserve context, but small enough to stay semantically focused. The right size depends on the use case, but the engineering principle is consistent: short chunks reduce noise and improve retrieval precision.

A basic chunker takes a document and a chunk size, then slices the text into contiguous blocks. Many systems also add overlap so that adjacent chunks share some context. Overlap is helpful when the relevant fact is split across boundaries, but too much overlap can create redundant retrieval results.

Deterministic chunking matters because it makes the retrieval system stable and easy to test. In this project, chunk size and overlap are explicit parameters, and invalid combinations are rejected. That prevents accidental behavior where a chunker silently produces empty or irregular segments.

Good chunking turns a document into a set of retrieval units. Each unit can be ranked independently against a query, which allows the system to identify the most relevant slice of information instead of retrieving an entire article.
9 changes: 9 additions & 0 deletions fixtures/documents/04-embeddings-and-vector-search.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Embeddings and Vector Search

Embeddings transform text into dense vectors so that similarity can be computed with mathematical operations instead of string comparisons. In a retrieval pipeline, each chunk receives an embedding that captures semantic relationships between words and phrases.

A mock embedding provider is useful in education because it is deterministic and reproducible. It does not require external models or API calls. Instead, the provider derives a stable vector from the chunk text using a hashing-based approach. This makes the workshop easy to run locally and ensures tests remain repeatable.

Once embeddings are created, a vector store such as FAISS can index the vectors. The store keeps both the vector and a reference to the chunk metadata, so the retriever can return the original text and source file. Query-time retrieval is then a nearest-neighbor search: convert the query into an embedding, compare it against the stored vectors, and select the best matches.

The ranking function is measured by distance. Lower distance means a stronger match, and a retrieval layer can convert that to a score that is easier to explain in a result table.
9 changes: 9 additions & 0 deletions fixtures/documents/05-mcp.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# MCP Tool Contracts

The Model Context Protocol defines a clean contract for connecting clients and servers. In practice, an MCP server exposes tools, resources, and prompts that other systems can inspect and invoke programmatically. This makes tool access predictable, structured, and machine-readable.

A tool contract includes a name, a description, and input schema. The client can ask the server which tools are available before invoking one. That allows an agent to discover capabilities without hard-coding them in advance. In a retrieval workflow, a tool may accept a user query and return the top-k most relevant chunks.

The retrieval tool in this project is intentionally thin. It validates the input query, invokes the retriever, and converts the returned results into a JSON-friendly payload. The MCP boundary is separated from the retrieval logic so that the same retrieval behavior can be reused outside a protocol implementation.

This explicit boundary is helpful for debugging. The protocol defines the interface, while the underlying code still owns the actual search semantics.
9 changes: 9 additions & 0 deletions fixtures/documents/06-evaluation-and-workflows.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Evaluation and Retrieval Workflows

A robust RAG workflow is not just about retrieving a result. It is also about evaluating whether the right evidence was selected and whether the final response is grounded in that evidence. Evaluation can involve checking source coverage, measuring retrieval precision, and reviewing the final answer for unsupported claims.

In a workshop, a small retrieval pipeline often uses hand-crafted queries to test the system. For example, a user may ask, "What is retrieval augmented generation?" or "What does an MCP tool do?" The expected behavior is that the retrieval layer surfaces the most relevant passages rather than a random assortment of semantically adjacent text.

A good workflow keeps the retrieval stage deterministic. That makes it easy to compare the effects of different chunk sizes, overlap values, and top-k limits. It also reduces the chance that small changes in the environment will change the results in surprising ways.

This project emphasizes a transparent, local workflow that is easy to reason about. The documents, embeddings, and retrieval results are all inspectable, which makes the lab suitable for teaching and deliberate experimentation.
3 changes: 0 additions & 3 deletions fixtures/documents/mcp_basics.txt

This file was deleted.

3 changes: 0 additions & 3 deletions fixtures/documents/mcp_tools.txt

This file was deleted.

3 changes: 0 additions & 3 deletions fixtures/documents/rag_basics.txt

This file was deleted.

4 changes: 3 additions & 1 deletion src/rag/loaders.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,13 +37,15 @@ def load_documents(directory: str | Path) -> list[Document]:
continue

content = file_path.read_text(encoding="utf-8")
relative_source = file_path.relative_to(path).as_posix()
documents.append(
Document(
document_id=file_path.stem,
text=content,
metadata={
"source": str(file_path),
"source": relative_source,
"filename": file_path.name,
"document_type": suffix,
},
)
)
Expand Down
22 changes: 22 additions & 0 deletions tests/test_loaders.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,3 +40,25 @@ def test_load_documents_utf8_and_unsupported_extensions(tmp_path: Path) -> None:

assert [doc.filename for doc in documents] == ["hello.txt", "notes.md"]
assert documents[0].text == "héllo\n"
assert documents[0].source == "hello.txt"
assert "/tmp" not in documents[0].source


def test_load_documents_markdown_files_are_supported(tmp_path: Path) -> None:
doc_dir = tmp_path / "workshop"
doc_dir.mkdir()
(doc_dir / "overview.md").write_text("# Overview\n\nRAG helps retrieval.\n", encoding="utf-8")
(doc_dir / "notes.txt").write_text("plain text\n", encoding="utf-8")

documents = load_documents(doc_dir)

assert [doc.filename for doc in documents] == ["notes.txt", "overview.md"]
assert all(doc.source.endswith((".txt", ".md")) for doc in documents)


def test_workshop_corpus_is_a_markdown_collection() -> None:
documents = load_documents(Path(__file__).resolve().parents[1] / "fixtures" / "documents")

assert len(documents) == 6
assert all(doc.filename.endswith(".md") for doc in documents)
assert all(doc.metadata["source"].endswith(".md") for doc in documents)
Loading