paper-extract turns scientific-paper PDFs into a Markdown library with extracted figures and tables. It can also generate a compact, structured card for each paper so later prompts can cite the card instead of including the full paper.
marker performs PDF conversion. Optional conversion cleanup and card generation use models through OpenRouter.
- Python 3.13
- uv
- An OpenRouter API key for card generation or conversion cleanup
llama.cppon Apple Silicon when marker uses its vision model- Zotero 10 with its local API enabled and the optional
zoteroextra for tagged-paper discovery
brew install llama.cppThe first conversion downloads several gigabytes of marker model weights to ~/.cache/huggingface. Later runs reuse them.
Install the base command directly from GitHub:
uv tool install git+https://github.com/adbX/paper-extract.git
paper-extract --helpThe base installation supports manual PDF conversion, catalog indexing, and offline search. Tagged Zotero discovery additionally requires paper-extract[zotero]:
uv tool install 'paper-extract[zotero] @ git+https://github.com/adbX/paper-extract.git'The extra requires ZotMD 0.5. Until that release is published, clone zotmd beside this repository and use the editable development setup:
git clone https://github.com/adbX/zotmd.git ../zotmd
git clone https://github.com/adbX/paper-extract.git
cd paper-extract
uv sync --extra zotero
uv run paper-extract --helpSet OPENROUTER_API_KEY in the process environment or in a .env file. Credentials cannot be stored in the TOML configuration.
Without --config, .env is loaded from the invocation directory. With --config, it is loaded from the configuration file's directory. An existing process-level value takes precedence over .env.
OPENROUTER_API_KEY=sk-or-...card exits if no key is available. sync also exits before conversion when cards are missing, unless --no-card is supplied. convert warns once and disables its optional LLM cleanup pass. index, search, list, and results require no key and make no network requests.
The default workspace uses papers/pdfs for manual input, papers/md for output, and paper-library.jsonl for the catalog, relative to the invocation directory:
paper-extract sync --dry-run
paper-extract sync
paper-extract search '10.1000/example' --field doi --json
paper-extract list --format jsonlsync reconciles manual PDFs and, when [zotero] is configured, tagged Zotero papers. It converts new sources, generates missing cards, harvests missing results, and rebuilds the catalog. The dry run performs source discovery and reports conversion, paid card, results, duplicate, stale, unresolved, and inactive counts without writing. Use --no-card to avoid paid card generation, --no-results to skip result harvesting, and --refresh-changed to transactionally replace stale output.
index deliberately reconciles local per-paper state and recognizable legacy output, then replaces the catalog. sync is the source-aware writer and also rebuilds the catalog after completing its pipeline. search and list strictly load and validate the existing catalog without discovery or writes. A Git-pulled catalog is therefore a durable offline snapshot even when the consumer has no local *_paper.json state. A missing or malformed catalog is an error that directs the current writer to run paper-extract index or paper-extract sync.
search --field accepts doi, arxiv, citation-key, zotero-key, title, author, venue, tag, or year. Human output is the default; search --json returns ranked matches with state and resolved artifact paths. list --format jsonl emits resolved catalog records for agents, while list --format tsv emits directory, preview, title, authors, year, venue, identifiers, citation key, and state for external pickers.
The original stage-specific commands remain available:
paper-extract convert
paper-extract card
paper-extract resultsEach stage-specific command skips complete work that already exists. convert rejects a same-named output directory without matching paper state rather than treating partial or legacy output as complete. card --force regenerates existing cards. results --force rebuilds existing results/ folders, including any manual curation.
Run paper-extract <command> --help for all command options.
Pass a TOML file as a root option before the subcommand:
paper-extract --config paper-extract.toml convert
paper-extract --config paper-extract.toml card
paper-extract --config paper-extract.toml sync --dry-run
paper-extract --config paper-extract.toml search 'paper title'PAPER_EXTRACT_CONFIG can select the file when --config is absent. There is no implicit home-directory or parent-directory search. See paper-extract.example.toml for a complete generic example.
Only the following schema is accepted:
[paths]
pdfs = "papers/pdfs"
out = "papers/md"
catalog = "paper-library.jsonl"
[zotero]
tag = "paper-extract"
[convert]
mode = "fast" # "fast" or "balanced"
force_ocr = false
ocr_inline_math = true
highres_dpi = 192
use_llm = true
redo_inline_math = false
model = "google/gemini-3.5-flash"
[card]
model = "google/gemini-2.5-flash"
context = "research-focus.md"
# prompt = "card-prompt.md" # complete override; cannot be combined with contextUnknown sections and keys are rejected. Configured modes, value types, paths, models, and DPI are validated before marker or OpenAI is imported. force is deliberately CLI-only because persistent force settings could cause paid calls or erase curated output.
The presence of [zotero] enables exact manual-tag discovery during sync; omitting the table keeps every command independent of ZotMD. Paper-extract stores only this selection tag. ZotMD owns local API access and Zotero record parsing.
Precedence is:
--configoverPAPER_EXTRACT_CONFIGwhen selecting a file.- Explicit subcommand flags over TOML values.
- TOML values over built-in defaults.
- A process-level
OPENROUTER_API_KEYover the selected.envfile.
Relative paths in TOML resolve from the TOML file's directory. Built-in paper paths also resolve there when a config is selected, so that directory becomes the workspace root. Explicit CLI paths remain relative to the invocation directory.
Configurable booleans have positive and negative CLI forms. For example, --ocr-inline-math can override false in TOML, while --no-ocr-inline-math can override true.
This repository ships a cross-tool paper-library skill in .claude/skills/paper-library. The directory name identifies the shared source location; using the skill does not require Claude Code.
Keep a local clone so the skill can stay linked to the version you update. Install the zotero extra when the catalog should discover tagged Zotero papers, then copy and edit the example configuration:
git clone https://github.com/adbX/paper-extract.git
cd paper-extract
uv tool install 'paper-extract[zotero] @ git+https://github.com/adbX/paper-extract.git'
mkdir -p "$HOME/.config/paper-extract"
cp paper-extract.example.toml "$HOME/.config/paper-extract/paper-extract.toml"
export PAPER_EXTRACT_CONFIG="$HOME/.config/paper-extract/paper-extract.toml"Set PAPER_EXTRACT_CONFIG in the environment that launches the agent. Paths in the copied TOML remain relative to its directory until you replace them. Run paper-extract sync --dry-run to check source discovery and paid card counts, then paper-extract sync or paper-extract index to create the catalog.
OpenCode scans the Claude-compatible skill root, while Codex and OMP use the shared agents root. From the repository clone, deploy one direct link and one shared-root link:
mkdir -p "$HOME/.claude/skills" "$HOME/.agents/skills"
ln -s "$(pwd -P)/.claude/skills/paper-library" "$HOME/.claude/skills/paper-library"
ln -s "$HOME/.claude/skills/paper-library" "$HOME/.agents/skills/paper-library"Add the following rule to the relevant project or global AGENTS.md:
## Scholarly papers
Before web search or PDF retrieval for a scholarly paper, technical article, DOI, arXiv ID, or citation-key lookup, use the `paper-library` skill and search the configured local catalog. Fall back normally when `paper-extract`, `PAPER_EXTRACT_CONFIG`, or a usable local match is absent.The packaged prompt works without configuration and produces a generic card. To retain its schema while tailoring the final synthesis section, provide a UTF-8 research-context file:
paper-extract card --context research-focus.mdThe file's text is inserted literally. Braces, Markdown fences, and pagination-shaped text are not interpreted.
To replace the complete system prompt, use:
paper-extract card --prompt card-prompt.md--prompt and --context are mutually exclusive. Their TOML equivalents are [card].prompt and [card].context. Files must be regular, non-empty UTF-8 files. A full override is read unchanged and receives no interpolation or appended context.
sync discovers papers through Zotero 10's local HTTP API and reads validated Zotero-managed PDF files in place. Paper-extract does not send discovery requests or PDFs through Zotero's service. Zotero's own configured data synchronization and WebDAV file synchronization remain outside paper-extract, and the separate zotmd sync command continues to use Zotero's Web API.
PDF conversion runs marker locally. When conversion cleanup is enabled, marker can send selected text blocks or page crops to the configured OpenRouter model. Card generation sends the complete converted paper Markdown to the selected model. Result harvesting, catalog indexing, search, and listing remain local. Review sync --dry-run before a batch to see the number of paid card operations.
paper-library.jsonl
papers/
├── pdfs/
└── md/
└── <paper>/
├── <paper>.md
├── <paper>-card-<model>.md
├── <paper>_meta.json
├── <paper>_paper.json
└── results/
Converted Markdown starts with deterministic producer-owned bibliographic frontmatter. The ignored <paper>_paper.json file records source identity and fingerprint, conversion settings, artifact paths, and stale or unresolved state. Absolute source paths remain confined to that ignored state file. The deterministic JSONL catalog stores only workspace-relative artifact paths and bibliographic metadata; query output resolves those paths at runtime.
New conversions are built in a temporary sibling directory and become visible only after Markdown, metadata, images, frontmatter, and state are complete. Explicit refresh operations use the same staging boundary for conversion, card generation, and result harvesting, then replace the old directory only after every requested phase succeeds.
The public repository ignores everything under papers/, the default generated catalog, local configuration, prompt and context files, per-paper state, and interrupted transaction files because they are user data or can contain local paths.
uv sync
uv run pytestFor local Zotero integration development, keep the editable ZotMD checkout at ../zotmd and include the extra:
uv sync --extra zotero
uv run --extra zotero paper-extract --config paper-extract.toml sync --dry-runApache-2.0. marker is licensed separately: its code uses Apache-2.0 and its model weights have their own licenses.