Scholar RAG Agent is a production-grade, local-first Agentic RAG system for scientific literature. It ingests papers from PDFs, arXiv, Semantic Scholar, OpenAlex, PubMed, PubMed Central (PMC), Crossref, Europe PMC, DOAJ, DBLP, HAL, OpenAIRE, Zenodo, Figshare, CORE, bioRxiv/medRxiv, NASA ADS, DataCite, OpenCitations, and OSF; builds hybrid dense, sparse, and entity-relationship retrieval indexes; and answers research questions with multi-hop reasoning and citation-backed evidence. Scholar RAG Agent is a production-grade, local-first Agentic RAG system for scientific literature. It ingests papers from PDFs, arXiv, Semantic Scholar, OpenAlex, PubMed, PubMed Central (PMC), Crossref, Europe PMC, DOAJ, DBLP, HAL, OpenAIRE, Zenodo, Figshare, CORE, bioRxiv/medRxiv, NASA ADS, DataCite, OpenCitations, and ORCID; builds hybrid dense, sparse, and entity-relationship retrieval indexes; and answers research questions with multi-hop reasoning and citation-backed evidence. Scholar RAG Agent is a production-grade, local-first Agentic RAG system for scientific literature. It ingests papers from PDFs, arXiv, Semantic Scholar, OpenAlex, PubMed, PubMed Central (PMC), Crossref, Europe PMC, DOAJ, DBLP, HAL, OpenAIRE, Zenodo, Figshare, CORE, bioRxiv/medRxiv, NASA ADS, DataCite, OpenCitations, OSF, and Unpaywall; builds hybrid dense, sparse, and entity-relationship retrieval indexes; and answers research questions with multi-hop reasoning and citation-backed evidence. Scholar RAG Agent is a production-grade, local-first Agentic RAG system for scientific literature. It ingests papers from PDFs, arXiv, Semantic Scholar search and recommendations, OpenAlex, PubMed, PubMed Central (PMC), Crossref, Europe PMC, DOAJ, DBLP, HAL, OpenAIRE, Zenodo, Figshare, CORE, bioRxiv/medRxiv, NASA ADS, DataCite, OpenCitations, OSF, and ORCID; builds hybrid dense, sparse, and entity-relationship retrieval indexes; and answers research questions with multi-hop reasoning and citation-backed evidence. Scholar RAG Agent is a production-grade, local-first Agentic RAG system for scientific literature. It ingests papers from PDFs, arXiv, Semantic Scholar, OpenAlex, PubMed, PubMed Central (PMC), Crossref, Europe PMC, DOAJ, DBLP, HAL, OpenAIRE, Zenodo, Figshare, CORE, bioRxiv/medRxiv, NASA ADS, DataCite, OpenCitations, OSF, ORCID, Unpaywall, and OpenAlex retraction alerts; builds hybrid dense, sparse, and entity-relationship retrieval indexes; and answers research questions with multi-hop reasoning and citation-backed evidence.
The project is designed for the scientific knowledge synthesis narrative behind NIW-style research impact: researchers can accelerate literature review, hypothesis validation, and grounded comparison across large corpora without losing provenance.
Most literature workflows break down when the corpus grows beyond a few papers:
-
Issue: keyword search misses papers that use different terminology. Scholar RAG Agent combines dense semantic retrieval, BM25 sparse search, HyDE expansion, and RRF fusion so a query can match both exact terms and related scientific phrasing.
-
Issue: fused results are dominated by near-duplicate passages that waste the context window. An optional Maximal Marginal Relevance (MMR) re-ranker balances relevance against novelty, dropping redundant chunks so the model sees complementary evidence.
-
Issue: single-hop RAG retrieves isolated snippets but misses evidence chains. The GraphRAG layer extracts entities and relationships, then follows bounded multi-hop paths to connect methods, datasets, findings, and limitations across papers.
-
Issue: generated summaries sound plausible but are hard to audit. Every answer is mapped back to retrieved chunk IDs, and unsupported claims are flagged with
[UNGROUNDED]instead of being silently trusted. -
Issue: research questions often need a plan, not just one search call. The Observe -> Decide -> Act runtime classifies intent, decomposes the query into retrieval sub-tasks, and persists a JSON rationale trace for every decision.
-
Issue: teams need reproducible evidence trails for reviews, grants, and publications. The SQLite event log records state transitions, timestamps, agent IDs, run IDs, plans, retrieval payloads, and final answer provenance.
-
Systematic literature review: ingest a folder of PDFs plus arXiv IDs, ask for the strongest themes, and receive cited claims grouped by supporting chunks.
-
Grant or NIW evidence synthesis: collect papers around a research contribution, validate novelty claims, and export citation-backed reasoning traces that show why each claim is supported.
-
Hypothesis validation: ask whether the literature supports or refutes a hypothesis, then inspect supporting and counter-evidence retrieval tasks separately.
-
Method comparison: compare approaches such as GraphRAG, dense retrieval, and BM25 across papers while preserving the source chunks behind each contrast.
-
Research onboarding: give a new lab member a paper corpus and let them ask grounded factual, synthesis, comparison, and hypothesis questions without manually reading every PDF first.
-
Prior-art triage: search Semantic Scholar and arXiv records, expand a trusted seed through Semantic Scholar recommendations, ingest abstracts, then identify overlapping methods, datasets, and claims before deeper manual review.
-
Citation QA for drafts: paste draft claims as questions and flag statements that are not supported by the ingested source chunks.
-
Multi-provider LLM evaluation: route reasoning, speed, cost, and default tasks to different adapters while keeping output validation and citation grounding consistent.
+---------------------------+
| Observe: Query Analyzer |
+-------------+-------------+
|
v
+---------+ +-------------+-------------+ +-------------------+
| Papers +----->| Decide: Planner +----->| Act: Executor |
+---------+ +-------------+-------------+ +---------+---------+
PDF/arXiv/S2 | |
v v
+-----------+-----------+ +----------+----------+
| SQLite Durable Events | | Hybrid Retrieval |
+-----------------------+ | Dense + BM25 + RRF |
+----------+----------+
|
v
+----------+----------+
| GraphRAG Multi-hop |
+----------+----------+
|
v
+----------+----------+
| LLM Router + Guard |
+----------+----------+
|
v
Citation-backed answer
git clone https://github.com/Francis1998/scholar-rag-agent.git
cd scholar-rag-agent && uv sync --extra dev
uv run pytest tests/ -vuv run python scripts/demo_local.py
uv run uvicorn api.main:app --reloadThe deterministic demo ingests a small fixture paper, executes an Observe -> Decide -> Act run, prints the planner trace, and returns a cited answer. A generated demo asset is available at docs/assets/demo.gif.
Additional GIFs in docs/assets/ show the problem-to-solution flow, planner trace, and citation grounding guard.
| Document | Description |
|---|---|
| Quickstart | Install, demo, and API in three steps. |
| Architecture | Agent state machine, retrieval pipeline, and data flow. |
| Configuration | Environment variables and provider keys. |
| Configuration (extended) | Full configuration reference with examples. |
| Safety | Timeout policy, scope bounds, cancellation, and hallucination guard design. |
| Demo | Demo GIFs and reproducible local demo commands. |
| Examples | Usage examples for ingestion, querying, and retrieval evaluation. |
| Performance | Performance tuning notes. |
| Troubleshooting | Common setup and runtime fixes. |
| Contributing | Development and PR workflow. |
| Security | Vulnerability reporting policy. |
| Changelog | Version history. |
| bioRxiv / medRxiv source guide | bioRxiv and medRxiv preprint connector. |
| NASA ADS source guide | NASA ADS astronomy/physics connector. |
| PMC source guide | PubMed Central full-text connector. |
| DataCite source guide | DataCite DOI registry connector. |
| OpenCitations source guide | OpenCitations DOI metadata and citation-count connector. |
| Semantic Scholar recommendations guide | Related-paper expansion from a seed Semantic Scholar id or DOI. |
| OSF source guide | Open Science Framework preprint and registration connector. |
| ORCID source guide | ORCID public record works connector. |
| Unpaywall source guide | Unpaywall DOI open-access landing/PDF lookup connector. |
| Retraction check guide | OpenAlex retracted-works alert connector. |
| CORE source guide | CORE open-access works connector. |
| Figshare source guide | Figshare research-output connector. |
All live providers are optional. Without keys the system uses deterministic fakes for tests and demos. Configure keys in .env or your shell:
export OPENAI_API_KEY=...
export ANTHROPIC_API_KEY=...
export GEMINI_API_KEY=...
export MOONSHOT_API_KEY=...
export UNPAYWALL_EMAIL=dev@example.orgWhen enabled, downstream synthesis can route through GPT-5.5, Claude Sonnet 4.6, Gemini 3.x, and Kimi K2 while deterministic connectors such as Unpaywall keep source lookup reproducible. For downstream synthesis and evaluation, the preferred frontier model families are GPT-5.5, Claude Sonnet 4.6, Gemini 3.x, and Kimi K2. Gemini 3.x, and Kimi K2 while deterministic connectors such as Unpaywall and retraction checks keep source lookup reproducible.
uv run ruff check . && uv run ruff format --check .
uv run mypy src/
uv run pytest tests/ -v --cov=src --cov-fail-under=70Apache-2.0. See LICENSE.



