Multi-engine document OCR with cascading fallback and quality audit.
socr orchestrates multiple OCR engines — calling each as a CLI subprocess, auditing output quality, and falling back to a different engine when results are poor. Each engine is a standalone CLI tool (gemini-ocr, deepseek-ocr, marker-ocr, etc.) that can also be used independently.
pip install socr
# With specific engine backends
pip install socr[gemini] # Google Gemini (cloud)
pip install socr[local] # DeepSeek + Nougat (local/free)
pip install socr[all] # All enginesEngines are installed separately because they have different dependencies (torch, cloud SDKs, etc.). Install only what you need.
# Process a PDF: cost-aware agentic routing, page-major (the only control loop).
# Processes + saves one page at a time (pages/NNN.md), resumable on re-run,
# byte-identical final output. `--agentic` is accepted for compatibility but is
# always on — it's a no-op.
socr paper.pdf
socr paper.pdf --strict-local # local-only (free), page-by-page
socr paper.pdf --cost-budget 0.05 # cap spend per document
# Interrupted a run? Just run it again — finished pages are skipped:
socr paper.pdf # resumes from the last saved page
# Choose the first rung of the cost ladder
socr paper.pdf --primary gemini
socr paper.pdf --save-figures
# Batch process a directory
socr batch ~/Papers/ -o ./results/
socr batch ~/Papers/ --dry-run # preview what would be processed
# Reproducibly rebuild a document from a manifest (no model calls)
socr replay output/paper/manifest.json -o paper.md
# Check which engines are available
socr enginessocr routes each page to an OCR engine, checks the result, and re-tries on a different engine when the result is poor. There is one control loop — agentic, cost-aware, page-major — and it is always on.
PDF → for each page, in order:
route → extract text → reconcile tables → place figures/charts → (equations)
→ verify → FLUSH pages/NNN.md to disk → next page
→ stitch fragments → Markdown (byte-identical to whole-doc assembly) + manifest
The engine is chosen dynamically by cost: try the cheapest available provider
first, let a judge (a vision model that looks at the page, or a heuristic
fallback) decide accept-or-escalate, and climb the cost ladder
(local → cheap cloud → premium cloud) only when the cheaper output is rejected.
Stops at the first accepted output, bounded by --cost-budget / --max-cost-per-page.
Each run records the winning provider + cost per page and writes a manifest
that socr replay can reconstruct with zero model calls.
Progressive, page-by-page processing. In agentic mode socr is page-major: it
finishes one page completely and writes it to disk immediately (pages/NNN.md
- a
pages/NNN.jsonstatus sidecar) before starting the next, rather than holding the whole document in memory and saving once at the end. Concretely:
- Crash-safe. A hang or crash on page 30 leaves pages 1–29 already on disk. The
final
<stem>.mdis the ordered stitch of the fragments and is byte-identical to the non-progressive whole-doc assembly. - Resume. Re-running a partially-processed document skips pages already finished (matched by a per-page run-fingerprint + input checksum) and reprocesses only the rest. A model/prompt/flag change invalidates the fingerprint and forces re-OCR; the skip is conservative — on any doubt the page is reprocessed, never silently reused.
- Clean halt on a wedged model. If the local VLM stops responding, socr flushes
the finished pages, marks the document
PARTIAL_SAVE_VLM_TIMEOUT, and stops cleanly instead of feeding more work into a stuck GPU. - Tables / figures / charts, per page. Born-digital table geometry is verified before paying for a VLM judge; figures are embedded inline within their page; chart/front-matter pages (vector charts that would otherwise become text word-salad) are saved as image assets with a note rather than transcribed.
- Equations (opt-in).
--detect-equationssaves model-free crop PNGs of display equations;--recover-clean-equationsadditionally reads each crop to LaTeX into a non-destructive sidecar (validated; bad LaTeX never replaces the native text/crop).
Each engine is a separate CLI binary. socr calls it as a subprocess, reads the
output markdown, and applies the quality pipeline. See docs/ARCHITECTURE.md for
the full design.
Routing is native text → local qwen → marker → gemini (paid cloud edge case). See
docs/MODELS.md for the full per-sub-task policy and the measured data behind it.
| Engine | Package | Type | Routing role |
|---|---|---|---|
| Qwen | qwen-ocr-cli |
Local (Ollama) | Workhorse VLM (qwen3-vl:30b-a3b-instruct; the former qwen3.5:cloud was retired 2026-09-25) |
| Gemini | gemini-ocr-cli |
Cloud | Edge-case escalation, ~$0.0002/page |
| Marker | marker-ocr-cli |
Local | Layout-aware fallback (Surya + Texify) |
| GLM | glm-ocr-cli |
Local | Fast local emergency fallback |
| Nougat | nougat-ocr-cli |
Local | Academic papers, Python <3.13 |
| Mistral | mistral-ocr-cli |
Cloud | Manual only (--primary mistral); dominated by Gemini |
| DeepSeek | deepseek-ocr-cli |
Local | Manual only (--primary deepseek); low quality |
Check availability:
$ socr engines
[+] gemini cloud, ~$0.0002/page
[+] marker local, layout-aware (Surya + Texify)
[+] mistral cloud, ~$0.001/page
[+] deepseek local via Ollama
[x] nougat local, academic papers
socr process <PDF> [OPTIONS]
-o, --output-dir PATH Output directory
--primary ENGINE Primary OCR engine (gemini, marker, deepseek, etc.)
--fallback ENGINE REMOVED - rejected with an error (GH-142).
No execution path reads it; escalation IS the
fallback. Use --primary for the first rung.
--no-audit REMOVED - rejected with an error (GH-139)
--no-judge-hard-pages REMOVED - rejected with an error (GH-142)
--no-native-first OCR every page (don't use native text for prose)
--save-figures Extract figure PNGs + inline image refs (no captions)
--describe-figures Also add VLM captions (opt-in, non-authoritative)
--timeout SECONDS Subprocess timeout
--profile NAME Load ~/.config/socr/{name}.yaml
--config PATH Custom YAML config file
-q, --quiet / -v, --verbose Output verbosity
--dry-run / --reprocess List-only / force reprocess
# Agentic cost-aware routing (page-major; progressive save + resume; the only
# control loop — always on; --agentic is accepted for compatibility as a no-op)
--agentic No-op; cost-aware routing is always on. Per page:
cheapest provider first, judge escalates, then
flush pages/NNN.md to disk before the next page.
Re-running resumes from the last finished page.
--strict-local Only local/free rungs (no paid cloud)
--judge-backend MODE auto | vlm | heuristic (default: auto)
--judge-model NAME VLM model for the judge (e.g. qwen2-vl:7b)
--max-cost-per-page USD Skip providers above this price (0 = no cap)
--cost-budget USD Stop escalating once doc spend hits this (0 = ∞)
--write-manifest Write a replayable manifest + blob cache
--detect-equations Detect display-equation regions, save crop PNGs (model-free)
--recover-clean-equations Also read equation crops to LaTeX into a sidecar (opt-in)
socr batch <DIR> [OPTIONS]
Same options as process, plus:
--limit N Process first N files
socr replay <MANIFEST> [-o OUT] Rebuild a document from cache (no model calls)
socr judge-benchmark <DATASET> Score the judge against labeled good/mangled pages
socr engines Show available engines
socr library [--config PATH] [--dry-run] [--rerun STEM | --promote STEM]
Process the papers library from its own config
Reads the library's config (default ~/papers/config.yaml) instead of taking
paths on the command line. Every path comes from the config: ~ is expanded,
relative paths resolve against root, and a missing key or a path that escapes
root is an error, never a default. The output.document.* names must match
what the pipeline writes ({stem}.md, figures, metadata.json) or the load
fails.
socr library --dry-run # list PDFs under input.pdf that have no text yet
socr library # process them into output.text/{stem}/
socr library --rerun STEM # re-process an existing paper into staging
socr library --promote STEM # archive the old copy, install the staged one
- A PDF whose text directory exists is never written into. A text directory
without its markdown is reported and skipped; use
--rerun. - New papers are processed into the staging directory first and moved into
output.textwith a no-replace rename only after the pipeline produced{stem}.md. A crash or failure leaves its leftovers in staging (reported, never overwritten) andoutput.textuntouched. A finished run with statuspartial/failedis installed (best available text) and listed as unverified. - One run at a time: an exclusive lock (
.library.lockinindex.dir) is held for the whole run; a second run refuses. --rerunwrites to the staging directory (optional top-level config keystaging; default<root>/.socr-staging) and the stem is recorded asawaiting_approvalin the manifest. A staged run is never overwritten.--promoterenames the old text directory to<archive.dir>/<stem>.<YYYY-MM-DD>(a numeric suffix avoids a clash) and the staged one into place. Nothing is deleted. It refuses when nothing is staged. A journal (.promote.journal.jsoninindex.dir) brackets the two renames; the next run finishes an interrupted promotion before doing anything else.- The config is refused if index file names collide (case-insensitively), if the pdf/text/index/archive/staging directories are equal or nested (compared after symlink resolution), or if two PDFs share a stem case-insensitively.
- The library must live on a local filesystem. File locking and atomic renames are
not guaranteed on iCloud or network mounts (
~/papersis local by policy). - Renames use the kernel no-replace primitive (macOS
renamex_np(RENAME_EXCL), Linuxrenameat2(RENAME_NOREPLACE)); only where that is unavailable does it fall back to check-then-rename. A promotion journal naming paths outside the configured text/staging/archive dirs (or a symlink) is refused and left untouched. - An unreadable
unverified.txtaborts the index refresh; it is never read as empty. - After each run (not
--dry-run) the index is rewritten atomically:documents(absolute PDF paths),missing_text(stems),unverifiedandmanifest(per-stemstate:verified,unverifiedorunknown). A document isunverifiedonly on evidence: an explicit non-completedmetadata status, a pagewarning/error, or a hand-placedUNVERIFIED.txtin its text dir (read, never written or deleted). Legacy metadata with nostatusisunknownand is not listed.unverified.txtis the union of its existing entries and the computed ones; a stem leaves it only when this run processed it and it came outverified. backup.rclone_remoteis never read or written. The summary ends with "Run backup-gdrive to push".- Exit code is nonzero if any processed document failed or was partial.
output/<doc_stem>/
├── <doc_stem>.md # final text, stitched from pages/
├── metadata.json # document status and notes
├── pages/ # NNNNN.md body + NNNNN.json sidecar per page (resume ledger)
├── manifest.json # replay record; blobs in cache/
├── cache/ # content-addressed blobs the manifest points to
├── audit_log.json # notable events of the run
├── tables_trust.json # pages with doubtful tables (absent = none)
├── figures/ # images the text links to
└── equations/ # equation crops
The pages/ directory makes a run crash-safe and resumable: each NNNNN.md is
written the instant its page finishes. A re-run reuses a page only when its
NNNNN.json is terminal, the run fingerprint and input checksum match, the
status is success with audit_passed true, and the .md fragment is readable.
Three adjudicated outcomes are also reused although they are warning: a table the
ladder rejected (table_rejected) or withheld (table_withheld), and a page accepted on
a credentialed judge-timeout ladder (judge_timeout_ladder_accepted); see
_load_terminal_page in pipeline/orchestrator.py. Otherwise the page is reprocessed.
To judge whether a page can be trusted, start with status, failure_mode and
audit_passed in its pages/NNNNN.json, then read its audit events. Every file, status, failure mode and audit event is
explained in docs/OUTPUT.md.
Create ~/.config/socr/config.yaml:
primary_engine: gemini
fallback_engine: marker
timeout: 300
save_figures: false
audit_min_words: 50Every PipelineConfig field can be set here (see socr/core/config.py); keys under
hpc: are HPCConfig fields.
An unrecognised key is a hard error — the load fails and names the offending key.
This is deliberate: a silently-ignored config key is the bug this rule exists to prevent
(#240), where a cost_budget spend cap set in a file simply never took effect and
nothing said so.
Or use profiles: ~/.config/socr/fast.yaml → socr paper.pdf --profile fast
Each backend is an independent CLI tool:
- gemini-ocr-cli — Google Gemini
- deepseek-ocr-cli — DeepSeek via Ollama
- mistral-ocr-cli — Mistral AI
- marker-ocr-cli — Marker (Surya + Texify)
- nougat-ocr-cli — Meta Nougat
MIT