Skip to content

Repository files navigation

socr

PyPI Python 3.11–3.12 License

Multi-engine document OCR with cascading fallback and quality audit.

socr orchestrates multiple OCR engines — calling each as a CLI subprocess, auditing output quality, and falling back to a different engine when results are poor. Each engine is a standalone CLI tool (gemini-ocr, deepseek-ocr, marker-ocr, etc.) that can also be used independently.

Install

pip install socr

# With specific engine backends
pip install socr[gemini]          # Google Gemini (cloud)
pip install socr[local]           # DeepSeek + Nougat (local/free)
pip install socr[all]             # All engines

Engines are installed separately because they have different dependencies (torch, cloud SDKs, etc.). Install only what you need.

Usage

# Process a PDF: cost-aware agentic routing, page-major (the only control loop).
# Processes + saves one page at a time (pages/NNN.md), resumable on re-run,
# byte-identical final output. `--agentic` is accepted for compatibility but is
# always on — it's a no-op.
socr paper.pdf
socr paper.pdf --strict-local           # local-only (free), page-by-page
socr paper.pdf --cost-budget 0.05       # cap spend per document
# Interrupted a run? Just run it again — finished pages are skipped:
socr paper.pdf                          # resumes from the last saved page

# Choose the first rung of the cost ladder
socr paper.pdf --primary gemini
socr paper.pdf --save-figures

# Batch process a directory
socr batch ~/Papers/ -o ./results/
socr batch ~/Papers/ --dry-run        # preview what would be processed

# Reproducibly rebuild a document from a manifest (no model calls)
socr replay output/paper/manifest.json -o paper.md

# Check which engines are available
socr engines

How it works

socr routes each page to an OCR engine, checks the result, and re-tries on a different engine when the result is poor. There is one control loop — agentic, cost-aware, page-major — and it is always on.

Agentic, cost-aware routing (the only control loop)

PDF → for each page, in order:
        route → extract text → reconcile tables → place figures/charts → (equations)
        → verify → FLUSH pages/NNN.md to disk → next page
    → stitch fragments → Markdown (byte-identical to whole-doc assembly) + manifest

The engine is chosen dynamically by cost: try the cheapest available provider first, let a judge (a vision model that looks at the page, or a heuristic fallback) decide accept-or-escalate, and climb the cost ladder (local → cheap cloud → premium cloud) only when the cheaper output is rejected. Stops at the first accepted output, bounded by --cost-budget / --max-cost-per-page. Each run records the winning provider + cost per page and writes a manifest that socr replay can reconstruct with zero model calls.

Progressive, page-by-page processing. In agentic mode socr is page-major: it finishes one page completely and writes it to disk immediately (pages/NNN.md

  • a pages/NNN.json status sidecar) before starting the next, rather than holding the whole document in memory and saving once at the end. Concretely:
  • Crash-safe. A hang or crash on page 30 leaves pages 1–29 already on disk. The final <stem>.md is the ordered stitch of the fragments and is byte-identical to the non-progressive whole-doc assembly.
  • Resume. Re-running a partially-processed document skips pages already finished (matched by a per-page run-fingerprint + input checksum) and reprocesses only the rest. A model/prompt/flag change invalidates the fingerprint and forces re-OCR; the skip is conservative — on any doubt the page is reprocessed, never silently reused.
  • Clean halt on a wedged model. If the local VLM stops responding, socr flushes the finished pages, marks the document PARTIAL_SAVE_VLM_TIMEOUT, and stops cleanly instead of feeding more work into a stuck GPU.
  • Tables / figures / charts, per page. Born-digital table geometry is verified before paying for a VLM judge; figures are embedded inline within their page; chart/front-matter pages (vector charts that would otherwise become text word-salad) are saved as image assets with a note rather than transcribed.
  • Equations (opt-in). --detect-equations saves model-free crop PNGs of display equations; --recover-clean-equations additionally reads each crop to LaTeX into a non-destructive sidecar (validated; bad LaTeX never replaces the native text/crop).

Each engine is a separate CLI binary. socr calls it as a subprocess, reads the output markdown, and applies the quality pipeline. See docs/ARCHITECTURE.md for the full design.

Engines

Routing is native text → local qwen → marker → gemini (paid cloud edge case). See docs/MODELS.md for the full per-sub-task policy and the measured data behind it.

Engine Package Type Routing role
Qwen qwen-ocr-cli Local (Ollama) Workhorse VLM (qwen3-vl:30b-a3b-instruct; the former qwen3.5:cloud was retired 2026-09-25)
Gemini gemini-ocr-cli Cloud Edge-case escalation, ~$0.0002/page
Marker marker-ocr-cli Local Layout-aware fallback (Surya + Texify)
GLM glm-ocr-cli Local Fast local emergency fallback
Nougat nougat-ocr-cli Local Academic papers, Python <3.13
Mistral mistral-ocr-cli Cloud Manual only (--primary mistral); dominated by Gemini
DeepSeek deepseek-ocr-cli Local Manual only (--primary deepseek); low quality

Check availability:

$ socr engines

  [+] gemini       cloud, ~$0.0002/page
  [+] marker       local, layout-aware (Surya + Texify)
  [+] mistral      cloud, ~$0.001/page
  [+] deepseek     local via Ollama
  [x] nougat       local, academic papers

CLI reference

socr process <PDF> [OPTIONS]
  -o, --output-dir PATH       Output directory
  --primary ENGINE             Primary OCR engine (gemini, marker, deepseek, etc.)
  --fallback ENGINE            REMOVED - rejected with an error (GH-142).
                               No execution path reads it; escalation IS the
                               fallback. Use --primary for the first rung.
  --no-audit                   REMOVED - rejected with an error (GH-139)
  --no-judge-hard-pages        REMOVED - rejected with an error (GH-142)
  --no-native-first            OCR every page (don't use native text for prose)
  --save-figures               Extract figure PNGs + inline image refs (no captions)
  --describe-figures           Also add VLM captions (opt-in, non-authoritative)
  --timeout SECONDS            Subprocess timeout
  --profile NAME               Load ~/.config/socr/{name}.yaml
  --config PATH                Custom YAML config file
  -q, --quiet / -v, --verbose  Output verbosity
  --dry-run / --reprocess      List-only / force reprocess

  # Agentic cost-aware routing (page-major; progressive save + resume; the only
  # control loop — always on; --agentic is accepted for compatibility as a no-op)
  --agentic                    No-op; cost-aware routing is always on. Per page:
                               cheapest provider first, judge escalates, then
                               flush pages/NNN.md to disk before the next page.
                               Re-running resumes from the last finished page.
  --strict-local               Only local/free rungs (no paid cloud)
  --judge-backend MODE         auto | vlm | heuristic (default: auto)
  --judge-model NAME           VLM model for the judge (e.g. qwen2-vl:7b)
  --max-cost-per-page USD      Skip providers above this price (0 = no cap)
  --cost-budget USD            Stop escalating once doc spend hits this (0 = ∞)
  --write-manifest             Write a replayable manifest + blob cache
  --detect-equations           Detect display-equation regions, save crop PNGs (model-free)
  --recover-clean-equations    Also read equation crops to LaTeX into a sidecar (opt-in)

socr batch <DIR> [OPTIONS]
  Same options as process, plus:
  --limit N                    Process first N files

socr replay <MANIFEST> [-o OUT]  Rebuild a document from cache (no model calls)
socr judge-benchmark <DATASET>   Score the judge against labeled good/mangled pages
socr engines                     Show available engines
socr library [--config PATH] [--dry-run] [--rerun STEM | --promote STEM]
                                 Process the papers library from its own config

Papers library (socr library)

Reads the library's config (default ~/papers/config.yaml) instead of taking paths on the command line. Every path comes from the config: ~ is expanded, relative paths resolve against root, and a missing key or a path that escapes root is an error, never a default. The output.document.* names must match what the pipeline writes ({stem}.md, figures, metadata.json) or the load fails.

socr library --dry-run             # list PDFs under input.pdf that have no text yet
socr library                       # process them into output.text/{stem}/
socr library --rerun STEM          # re-process an existing paper into staging
socr library --promote STEM        # archive the old copy, install the staged one
  • A PDF whose text directory exists is never written into. A text directory without its markdown is reported and skipped; use --rerun.
  • New papers are processed into the staging directory first and moved into output.text with a no-replace rename only after the pipeline produced {stem}.md. A crash or failure leaves its leftovers in staging (reported, never overwritten) and output.text untouched. A finished run with status partial/failed is installed (best available text) and listed as unverified.
  • One run at a time: an exclusive lock (.library.lock in index.dir) is held for the whole run; a second run refuses.
  • --rerun writes to the staging directory (optional top-level config key staging; default <root>/.socr-staging) and the stem is recorded as awaiting_approval in the manifest. A staged run is never overwritten.
  • --promote renames the old text directory to <archive.dir>/<stem>.<YYYY-MM-DD> (a numeric suffix avoids a clash) and the staged one into place. Nothing is deleted. It refuses when nothing is staged. A journal (.promote.journal.json in index.dir) brackets the two renames; the next run finishes an interrupted promotion before doing anything else.
  • The config is refused if index file names collide (case-insensitively), if the pdf/text/index/archive/staging directories are equal or nested (compared after symlink resolution), or if two PDFs share a stem case-insensitively.
  • The library must live on a local filesystem. File locking and atomic renames are not guaranteed on iCloud or network mounts (~/papers is local by policy).
  • Renames use the kernel no-replace primitive (macOS renamex_np(RENAME_EXCL), Linux renameat2(RENAME_NOREPLACE)); only where that is unavailable does it fall back to check-then-rename. A promotion journal naming paths outside the configured text/staging/archive dirs (or a symlink) is refused and left untouched.
  • An unreadable unverified.txt aborts the index refresh; it is never read as empty.
  • After each run (not --dry-run) the index is rewritten atomically: documents (absolute PDF paths), missing_text (stems), unverified and manifest (per-stem state: verified, unverified or unknown). A document is unverified only on evidence: an explicit non-completed metadata status, a page warning/error, or a hand-placed UNVERIFIED.txt in its text dir (read, never written or deleted). Legacy metadata with no status is unknown and is not listed. unverified.txt is the union of its existing entries and the computed ones; a stem leaves it only when this run processed it and it came out verified.
  • backup.rclone_remote is never read or written. The summary ends with "Run backup-gdrive to push".
  • Exit code is nonzero if any processed document failed or was partial.

Output

output/<doc_stem>/
├── <doc_stem>.md        # final text, stitched from pages/
├── metadata.json        # document status and notes
├── pages/               # NNNNN.md body + NNNNN.json sidecar per page (resume ledger)
├── manifest.json        # replay record; blobs in cache/
├── cache/               # content-addressed blobs the manifest points to
├── audit_log.json       # notable events of the run
├── tables_trust.json    # pages with doubtful tables (absent = none)
├── figures/             # images the text links to
└── equations/           # equation crops

The pages/ directory makes a run crash-safe and resumable: each NNNNN.md is written the instant its page finishes. A re-run reuses a page only when its NNNNN.json is terminal, the run fingerprint and input checksum match, the status is success with audit_passed true, and the .md fragment is readable. Three adjudicated outcomes are also reused although they are warning: a table the ladder rejected (table_rejected) or withheld (table_withheld), and a page accepted on a credentialed judge-timeout ladder (judge_timeout_ladder_accepted); see _load_terminal_page in pipeline/orchestrator.py. Otherwise the page is reprocessed.

To judge whether a page can be trusted, start with status, failure_mode and audit_passed in its pages/NNNNN.json, then read its audit events. Every file, status, failure mode and audit event is explained in docs/OUTPUT.md.

Configuration

Create ~/.config/socr/config.yaml:

primary_engine: gemini
fallback_engine: marker
timeout: 300
save_figures: false
audit_min_words: 50

Every PipelineConfig field can be set here (see socr/core/config.py); keys under hpc: are HPCConfig fields.

An unrecognised key is a hard error — the load fails and names the offending key. This is deliberate: a silently-ignored config key is the bug this rule exists to prevent (#240), where a cost_budget spend cap set in a file simply never took effect and nothing said so.

Or use profiles: ~/.config/socr/fast.yaml → socr paper.pdf --profile fast

Engine CLIs

Each backend is an independent CLI tool:

License

MIT

About

Multi-engine OCR with cascading fallback, quality audit, and figure extraction

Topics

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages