Run DINOv2 vision models in pure C++ on ggml. No Python, no PyTorch, no system dependencies at runtime.
See the GitHub releases page for the latest prebuilt binaries.
Release post & benchmarks → https://alexlavaee.me/projects/dinov2cpp/
Try it in your browser: the wasm demo runs the encoder fully client-side, nothing to install.
Using Claude Code, Cursor or another coding agent? Paste this:
Install https://github.com/espetro/dinov2.cpp via mise (latest), download the small GGUF from the dinov2-cpp-core profile on Hugging Face, and run classification plus `--print-embeddings` on an example image to verify it works.
Grab a prebuilt binary, download a weight, run one command. No toolchain needed.
1. Install a prebuilt binary from releases:
# Resolve the latest release tag (tarball names embed it)
TAG=$(basename "$(curl -sIL -o /dev/null -w '%{url_effective}' https://github.com/espetro/dinov2.cpp/releases/latest)")
# macOS (Apple Silicon)
curl -LO "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/espetro/dinov2.cpp/releases/download/$TAG/dinov2-$TAG-bin-macos-arm64.tar.gz"
tar xzf "dinov2-$TAG-bin-macos-arm64.tar.gz"
# Linux x64 / arm64: dinov2-$TAG-bin-ubuntu-x64.tar.gz / dinov2-$TAG-bin-ubuntu-arm64.tar.gz
# Windows (version-stable asset name):
# https://github.com/espetro/dinov2.cpp/releases/latest/download/dinov2-bin-win-cpu-x64.zip2. Download a GGUF weight:
hf download dinov2-cpp-core/dinov2-small-gguf --local-dir models3. Run inference (-c for classification; add -o out.png for a PCA feature visualization):
./bin/dinov2-cli -m models/model.gguf -i assets/tench.jpg -cEmbeddings (JSON) for PyTorch last_hidden_state[:, 0] consumers:
./bin/dinov2-cli -m models/model.gguf -i assets/tench.jpg --print-embeddingsEmits cls/pooled vectors on stdout as one JSON object.
Feature mode bounds the input's shortest edge to 518 px by default
(--preprocess bounded) so attention memory stays predictable on large
images; --preprocess hf mirrors the HF AutoImageProcessor recipe
(256 shortest edge + 224 center crop) for like-for-like comparisons, and
--preprocess crop518 fixes every input at 518x518 for batching mixed
aspect ratios. --no-resize keeps native resolution, and --max-tokens N
caps the patch-token count per image (default 4*(518/patch)^2).
Preview binary embeddings: for high-throughput consumers, use
--embeddings-binary -o output.d2e. This deliberately unstable preview writes
an explicit 40-byte little-endian header followed by float32 cls, pooled,
and optional row-major patch vectors. It writes no binary bytes to stdout.
With multiple inputs, -o is a directory containing indexed .d2e files.
The format may change without compatibility guarantees and is not a standard.
Batch inference: repeat -i (or comma-separate paths) and set --batch;
each input gets one JSON line on stdout, identical to running it alone
(see docs/cli.md):
./bin/dinov2-cli -m models/model.gguf -i a.jpg -i b.jpg --batch 2 --print-embeddingsThat's it. Committed benchmark rows, measured by CI on GitHub-hosted runners (dinov2-small, f16 GGUF, CPU, batch 1, forward pass only; full provenance in docs/benchmarks.md):
| Runner | Threads | Mean | Peak RSS |
|---|---|---|---|
ubuntu-latest (x86_64) |
12 | 255.0 ms | 104 MB |
ubuntu-24.04-arm (aarch64) |
4 | 133.8 ms | 96 MB |
For scale, an older desktop comparison on an i9-14900HX (24 threads) measured
small-with-registers at 64 ms vs 297 ms under PyTorch, with peak memory at
110 MB vs 455 MB. That table is retained as
historical context only and is not
reproducible from the CI workflow. Reproduce on your own hardware with
--bench --bench-json.
- Photo or asset dedup: examples/dedup/ clusters near-duplicates with the CLI and writes an HTML review page. Measured on a 360-image synthetic corpus: 124 s wall clock, precision 1.0, recall 0.977 at the default threshold.
- Embeddings over HTTP: tools/server/ is a single-binary
microservice; POST an image to
/v1/embeddings, get vectors back as JSON. - Running in the browser: the encoder compiles to WebAssembly (wasm/, docs/wasm.md); there is a live demo.
- CI visual gating: examples/ci-visual-regression/ is a copy-paste GitHub Action that compares screenshots by embedding cosine.
- Containers: examples/container/ builds a 45 MB scratch image.
- Library integration: embed
libdinov2through the C API in include/dinov2.h; no ggml types leak into the surface.
Not for: text or cross-modal search (embeddings are image-to-image only; use a CLIP-style model for text queries), HEIC/RAW/WebP/AVIF inputs without converting first, or corpora above ~1M images without a real ANN index (FAISS, sqlite-vec, ...) on top.
Embedding dinov2.cpp instead of shelling out? The build produces
libdinov2 (static by default, shared with -DBUILD_SHARED_LIBS=ON) and a
pure C header at include/dinov2.h: opaque handles, no
ggml types, status codes instead of aborts. Models load from a file, a memory
buffer, or a read callback; dino_encode runs a batch of raw RGB8 images and
the dino_output_* accessors return borrowed pointers into the context.
#include "dinov2.h"
dino_model *m = dino_model_load_from_file("models/model.gguf", dino_model_default_params());
dino_ctx *c = dino_init_from_model(m, dino_ctx_default_params());
dino_image img = {pixels, w, h, 0}; // interleaved RGB8, stride 0 = packed
if (dino_encode(c, &img, 1, dino_run_default_params()) == DINO_STATUS_SUCCESS) {
const float *cls = dino_output_cls(c, 0); // dino_model_hidden_size(m) floats
}
dino_free(c);
dino_model_free(m);The API is unstable for the 0.4.x line; see docs/stability.md.
Opt-in surfaces, off by default and not covered by the stability contract (docs/tiers.md):
- HTTP server:
dinov2-serveris a single-binary embeddings microservice over the C API (-DDINOV2_BUILD_SERVER=ON). POST an image to/v1/embeddings, getcls/pooled/patch vectors back as JSON. See tools/server/README.md. - wasm: the encoder compiles to WebAssembly and runs in the browser; see docs/wasm.md.
- examples/:
examples/dedup/finds near-duplicate photos with the CLI (stdlib-only Python, writes an HTML review page);examples/ci-visual-regression/is a copy-paste GitHub Action that gates screenshots by embedding cosine;examples/container/packages the CLI as a 45 MB scratch image.
| Feature | Detail |
|---|---|
| Zero dependencies | Image I/O via vendored stb, compute via ggml. Nothing else. |
| CPU, CUDA, Metal | Backends via ggml wherever ggml supports them. |
| f16 GGUF weights | Plus q4_0 through q8_0 quantization. |
| PyTorch-parity outputs | CLS + patch embeddings as JSON, matching the reference implementation. |
| Cross-platform prebuilts | macOS arm64, Linux x64/arm64, Windows x64. |
Limitations: image decode is vendored stb_image only (JPEG, PNG, BMP,
TGA; no HEIC, WebP, RAW, or AVIF, convert those first with sips,
heif-convert, or ffmpeg). Embeddings are image-to-image: this is not
CLIP and there are no text queries. Classification is ImageNet-1k only.
Ready-to-download f16 GGUF weights, published by CI to the dinov2-cpp-core Hugging Face profile:
| Model | GGUF download | Size | Recommendation |
|---|---|---|---|
| small (no registers) | dinov2-cpp-core/dinov2-small-gguf |
~50 MB | Baseline reproduction or task comparison |
| base (no registers) | dinov2-cpp-core/dinov2-base-gguf |
~180 MB | Baseline reproduction or task comparison |
| large (no registers) | dinov2-cpp-core/dinov2-large-gguf |
~620 MB | Baseline reproduction or task comparison |
| giant (no registers) | dinov2-cpp-core/dinov2-giant-gguf |
~2.2 GB | Baseline reproduction or task comparison |
| small (registers) | dinov2-cpp-core/dinov2-with-registers-small-gguf |
~50 MB | Recommended for patch/dense features |
| base (registers) | dinov2-cpp-core/dinov2-with-registers-base-gguf |
~180 MB | Recommended for patch/dense features |
| large (registers) | dinov2-cpp-core/dinov2-with-registers-large-gguf |
~620 MB | Recommended for patch/dense features |
| giant (registers) | dinov2-cpp-core/dinov2-with-registers-giant-gguf |
~2.2 GB | Recommended for patch/dense features |
Backbone-only checkpoints carry the same encoder weights without the ImageNet
classifier head. They run the feature modes (--print-embeddings,
--print-patch-tokens, -o PCA, --bench) and reject -c cleanly. The
dinov2-backbone-* repos below are wired into the
publish workflow but are not published on Hugging
Face yet; the table lists the target repo names and expected sizes:
| Model | GGUF repo (pending publish) | Size | Recommendation |
|---|---|---|---|
| small (backbone, no registers) | dinov2-cpp-core/dinov2-backbone-small-gguf |
~50 MB | Feature mode only; no classifier head |
| base (backbone, no registers) | dinov2-cpp-core/dinov2-backbone-base-gguf |
~180 MB | Feature mode only; no classifier head |
| large (backbone, no registers) | dinov2-cpp-core/dinov2-backbone-large-gguf |
~620 MB | Feature mode only; no classifier head |
| giant (backbone, no registers) | dinov2-cpp-core/dinov2-backbone-giant-gguf |
~2.2 GB | Feature mode only; no classifier head |
| small (backbone, registers) | dinov2-cpp-core/dinov2-backbone-with-registers-small-gguf |
~50 MB | Feature mode only; no classifier head |
| base (backbone, registers) | dinov2-cpp-core/dinov2-backbone-with-registers-base-gguf |
~180 MB | Feature mode only; no classifier head |
| large (backbone, registers) | dinov2-cpp-core/dinov2-backbone-with-registers-large-gguf |
~620 MB | Feature mode only; no classifier head |
| giant (backbone, registers) | dinov2-cpp-core/dinov2-backbone-with-registers-giant-gguf |
~2.2 GB | Feature mode only; no classifier head |
For patch and dense feature workflows, the register-token variants are the recommended starting point. In the settings studied in Vision Transformers Need Registers, register tokens reduce high-norm patch-token artifacts and produce smoother local feature and attention maps. Use them with --print-patch-tokens, PCA, dense features, and object discovery. This is not a universal accuracy claim: the official DINOv2 results show classification and retrieval results that depend on the task and model size. Choose the no-register variant for exact baseline reproduction or task-specific classification and retrieval comparisons.
- CONTRIBUTING.md: pre-converted GGUF weights, build from source, dev harness, tests, PR guidelines
- docs/cli.md:
dinov2-clireference: flags, output modes, embeddings JSON schema, workflows - docs/build.md: per-device optimizations, quantization
- docs/benchmarks.md: benchmarks against PyTorch, how to run your own
- docs/trust.md: the evidence behind the project's claims (nightly parity gate, sanitizer CI, benchmark provenance) and the limits of that evidence
- docs/hf-publishing.md: how CI publishes GGUF weights to Hugging Face
- docs/wasm.md: Emscripten/WebAssembly build and browser demo
- docs/stability.md + docs/tiers.md: the compatibility contract and tier policy
Upstream requires building from source and converting weights yourself. This fork ships what's missing: prebuilt binaries for macOS, Linux and Windows, plus ready-to-download GGUF weights published automatically by CI.
Forked from lavaman131/dinov2.cpp. Built on and heavily inspired by vit.cpp and ggml.

