Demos · Knowledge base

▁▂▃▅▃▂▁

Build a Knowledge Base

Set up a project, add a paper, check the copy, and close the session.

Before you start

Install Claude Code and the Open Science Skills (Getting started). The PDF converter also needs Java 11 or newer, and the check for scanned PDFs needs poppler-utils.

Whatever the agent reads is sent to your model provider. The sample files are invented. Before you use your own, check what your provider and your institution allow.

1. Set up the project

Open Claude Code in an empty folder:

mkdir my-project && cd my-project
claude
/oss:research-repo .
On the left, the folder layout the skill writes: originals gitignored, Markdown conversions tracked, a drop zone, a bibliography, the converter script, agent conventions, and a gitignore. On the right, the chain: a citation key in the manuscript looks up its bibliography entry, which matches by author and year to a tracked Markdown file, converted at intake from an original kept on local disk.
The folders the skill creates, and how a citation leads to its bibliography entry, the Markdown copy, and the original.

2. Add a paper

Put a PDF in sources/unprocessed/ and run:

/process-source

To use the four sample files instead:

git clone https://github.com/scdenney/ai-for-research.git ~/ai-for-research
cp ~/ai-for-research/demos/knowledge-base/inbox/* sources/unprocessed/

For each file, the agent:

PDFs are converted with OpenDataLoader PDF. Check the author and year the agent picks, because the file name and the key are built from them.

The intake pipeline in five steps: drop, identify, rename, convert, register. Under each step, what it leaves on disk, and what you or the agent do at that step. A branch off the convert step shows an image-only scan being reported as needing OCR instead of being converted.
One source, five steps.

3. Check the copy against the PDF

Open the Markdown copy next to the original.

On the left, the results table as printed in the original PDF: two outcomes, each with an estimate and a 95% interval in separate columns. On the right, the converted Markdown: the header row has joined the section heading, and both rows run together on one line, with all the numbers still present.
The results table in the sample Ferreira and Nair PDF, and the text the converter produced from it.

4. End the session

/oss:finished

This writes a handoff note (where things stand and what comes next) and adds an entry to the session log. The first time, it asks before creating these files. It does not commit. Read both, then commit:

git add -A
git status --short   # no line should start with sources/og/
git commit -m "Add the first sources"

The originals in sources/og/ are not committed, so back them up separately. If you push the project to GitHub, make the repository private: the copies are still copyrighted text.

5. Use it

The reference-check demo has a manuscript to run both checks on. The full walkthrough covers all four sample files and an audit of a messy project.