Literature-Driven LLM Generator for Biomedical Research
Oliver Bonham-Carter Β· with development testing by Vincent Mametjanov Email: obonhamcarter at allegheny.edu Β· GitHub
CuraLit extracts relevant articles from PubMed XML datasets and turns them into custom, literature-grounded AI assistants (via Ollama), a searchable statistics/visualization report, and a fact-verification database β so you can ask research questions and get answers backed by real, citable articles.
Looking for full command history and past release notes? See docs/CHANGELOG.md. The previous, more detailed README has been archived at docs/README_legacy_full.md.
- CuraLit π¬
CuraLit helps researchers:
- Search large PubMed datasets by keyword (streaming XML parser β scales to millions of articles)
- Analyze the matching corpus with statistics and interactive visualizations
- Verify facts with a local SQLite database (prevents AI hallucination of PMIDs/authors/DOIs)
- Retrieve grounded answers with RAG (Retrieval-Augmented Generation) β no fine-tuning required
- Generate a custom Ollama model fine-tuned on your curated literature (optional, alternative path)
All of this is available either from the command line, or from a local browser interface that exposes every command as a simple form.
Figure: interactive keyword/article network visualization.
Figure: frequency analysis of MeSH terms found in matching articles.
| Tool | Purpose | Required for |
|---|---|---|
| Rust 1.70+ | Build the curalit CLI |
Everything |
| Ollama | Run local LLMs and embeddings | RAG, model generation |
| Qdrant | Vector database | RAG only |
| uv | Python dependency management | Visualizations, browser UI |
git clone git@github.com:developmentAC/curalit.git
cd curalit
# Build the CLI (release build recommended)
cargo build --release
# Binary is at target/release/curalit
# Install Python dependencies (visualizations + browser UI)
uv syncDownload the PubMed .xml.gz files you want to search from the PubMed Baseline/Updatefiles FTP archives, extract them (gunzip), and place the resulting .xml files in data/.
CuraLit ships with a small local web app that exposes every command (search, stats, model generation, RAG build/query/ask, database build) as a form β no need to remember CLI flags.
uv run webui/app.py
# then open http://127.0.0.1:5050 in your browserThe web UI shells out to your compiled curalit binary, so anything you can do on the command line, you can do from the browser. Results, logs, and links to each run's Markdown report are shown directly on the page.
A screenshot of the search screen. Statistics can be created from the results.
A screenshot of the model-building page. Here Ollama models may be created to be used for chatting about the research.
A screenshot of the database construction page. A database is used in the project to ensure that the results remain factual and do not become distorted by the AI component of the project.
A screenshot of the screen where the user may interact with an Ollama model to brainstorm ideas, ask questions about methods, refine hypotheses and similar tasks which help to develop a research project.
All bash/CLI commands documented below continue to work exactly as before β the browser interface is an additional, optional way to drive the same functionality.
This walkthrough follows the same steps as sampleRunScripts/quickRunCommands_13_sept_vi.sh. Every command below can also be run from the browser interface.
curalit search -k "prosthetic" -k "joint" -k "analysis" -d ./data -o myResultsThis creates a single, simply-named run folder: 0_out/myResults/, containing:
results.csvβ the matched articlesstats.json/stats.logβ corpus statisticsvisualize.pyβ a ready-to-run visualization scriptreport.mdβ a human-readable Markdown summary of this run
uv run 0_out/myResults/visualize.pyOpens interactive HTML charts (keyword frequency, MeSH terms, journals, and a keyword-article network).
curalit db-build -k "prosthetic" -k "joint" -k "analysis" -d ./data -n myResultsCreates 0_out/myResults/database.db, a searchable SQLite database used to verify that any PMIDs/authors/DOIs mentioned by the AI actually exist in your corpus.
curalit rag-build -c 0_out/myResults/results.csvThis embeds your articles (via Ollama's nomic-embed-text model) into a local Qdrant vector store, so questions can be answered using the exact text of matching articles.
curalit rag-generate -m llama3 \
-q "Describe a research project concerning prosthetics. Provide several articles to read, comment on them, and give references." \
--use-db 0_out/myResults/database.dbThe answer, along with any verified citations, is printed to the terminal and appended to 0_out/curalit_articles_report.md (named after the RAG collection) so you keep a written record of every question you ask.
If you'd rather fine-tune a standalone Ollama model instead of (or in addition to) RAG:
curalit generate -c 0_out/myResults/results.csv -m my-medical-llm -b llama3
ollama create my-medical-llm -f 0_out/myResults/Modelfile_my-medical-llm
ollama run my-medical-llm| Command | Purpose |
|---|---|
curalit search |
Search PubMed XML files for matching articles |
curalit stats |
(Re)generate statistics/visualizations from a checkpoint CSV |
curalit generate |
Create an Ollama Modelfile + training data from a checkpoint |
curalit package |
Package model files into a distributable archive |
curalit db-build |
Build a SQLite fact-verification database |
curalit rag-build |
Build a RAG vector index from a checkpoint |
curalit rag-query |
Retrieve raw passages relevant to a question |
curalit rag-generate |
Generate a full, cited answer using RAG (+ optional DB verification) |
curalit rag-package |
Package a RAG vector index for distribution |
curalit big-help |
Print detailed in-terminal help with examples |
Run curalit <command> --help for the full list of flags for any command. See docs/DATABASE_FEATURE.md for detailed database/fact-verification documentation.
Every command writes into a single, simply-named run directory under 0_out/:
0_out/
myResults/
results.csv # matched articles (checkpoint)
stats.json # statistics (machine-readable)
stats.log # statistics (human-readable)
visualize.py # interactive visualization script
database.db # fact-verification database (if built)
Modelfile_<model> # Ollama configuration (if generated)
training_<model>.jsonl # fine-tuning data (if generated)
report.md # Markdown summary of everything run in this folder
If you run the same -o name twice, CuraLit will not silently overwrite your previous run β it creates name_2/, name_3/, etc. Use --resume to continue writing into the same folder instead.
RAG questions and answers (which aren't tied to a single search run) are appended to 0_out/<collection_name>_report.md so you have a running transcript of everything you've asked.
# Run the full Rust test suite
cargo test
# Run a specific test file
cargo test --test parser_test
# Integration tests requiring live services (Qdrant + Ollama)
cargo test --test rag_integration_test -- --ignoredSee tests/README.md for details on what each test suite covers.
"No XML files found" β confirm your -d/--data-dir points to a directory containing .xml files (not .xml.gz).
RAG commands fail to connect β make sure Qdrant is running (docker run -p 6333:6333 -p 6334:6334 -v $(pwd)/qdrant_storage:/qdrant/storage qdrant/qdrant) and Ollama has the embedding model installed (ollama pull nomic-embed-text).
Browser UI can't find the curalit binary β build it first with cargo build --release, or set CURALIT_BIN=/path/to/curalit before running uv run webui/app.py.
Too many articles matched β narrow your keywords or switch from OR to AND logic (-l and).
- docs/DATABASE_FEATURE.md β fact-verification database deep dive
- docs/QUICKSTART.md β condensed quick-start guide
- docs/CHANGELOG.md β release history
- docs/quarto/presentation β slide deck walkthrough
- tests/README.md β test suite overview
- docs/README_legacy_full.md β the previous, more exhaustive README
MIT β see opensource.org/licenses/MIT.
Oliver Bonham-Carter β obonhamcarter at allegheny.edu β oliverbonhamcarter.com
Check back often to see the evolution of the project!! This project is a work-in-progress. Updates will come periodically.
If you would like to contribute to this project, then please do! For instance, if you see some low-hanging fruit or task that you could easily complete, that could add value to the project, then I would love to have your insight.
Otherwise, please create an Issue for bugs or errors. Since I am a teaching faculty member at Allegheny College, I may not have all the time necessary to quickly fix the bugs. I welcome the OpenSource Community to further the development of this project. Much thanks in advance.
If you appreciate this project, please consider clicking the project's Star button. :-)
