Trafae is a book discovery service: it searches multiple open book catalogs at once, deduplicates the results, and fuses them into a single ranked list. It contains a Go API server and a React frontend.
GET /books/search fans a query out to every configured provider in parallel
and merges the responses with Reciprocal Rank Fusion. Duplicate books across
providers are collapsed into one result that keeps the per-source metrics.
Every parameter is optional: search by topic (topic, repeatable), browse by
genre alone (the genre doubles as the search term across providers when no
topic is given), or combine both with the filters below.
Paging is stateless: pass page (1-based, default 1) alongside limit (page
size). A page-N request fetches the requested page plus a few pages of
lookahead from every provider, fuses and sorts that pool, slices it, and
caches the pool for books.cache_ttl_in_second — page navigation within the
cached pool is served from memory in milliseconds with zero upstream calls,
and ordering stays stable across pages. Requests deeper than the cached pool
re-fetch with fresh lookahead. Providers cap their fetch depth (Open Library,
DOAB and loc.gov at 100 results; Internet Archive, Gutendex and Project
Gutenberg at 200), so very deep pages run out of pool and has_more honestly
turns false — with the frontend's limit of 24 that is around page 10.
| Query param | Meaning |
|---|---|
topic |
search term, repeatable (max 8); optional |
genre |
genre filter; defaults to default_genre, works standalone |
page |
1-based page number (max 20) |
limit |
page size (default 20, max 50) |
provider |
restrict to specific providers, repeatable |
language |
two-letter language filter |
min_year / max_year |
publication year range |
min_popularity |
provider-specific popularity (e.g. downloads) |
min_rating |
0–5 rating floor |
Supported providers:
| Provider | Popularity | Ratings | Notes |
|---|---|---|---|
| Open Library | — | 0–5 | Rate limited; set open_library_contact_email for higher limits |
| DOAB | — | — | Scholarly open-access books via the DSpace 7 API; links go to the DOAB record |
| Gutendex | Downloads | — | Project Gutenberg via Gutendex |
| Internet Archive | Downloads | 0–5 | |
| Library of Congress | — | — | Digital collections |
| Wikidata | — | — | Community-supplied metadata via the Wikidata Action API |
| Project Gutenberg | Downloads | — | Official catalog synced daily into SQLite |
GET /books/providers lists each provider with the filters it supports, so
clients can adapt queries to what will actually be served.
Examples:
curl "http://localhost:8081/books/search?topic=biology&genre=non-fiction&min_rating=4&limit=10"
curl "http://localhost:8081/books/search?genre=history&page=2&limit=24"make runThe API listens on http://localhost:8081. Health, readiness, and funnel
counters are available at /health, /ready, and /metrics.
On startup the server syncs the official Project Gutenberg catalog (a gzipped CSV) into SQLite and refreshes it every 24 hours. Searches against it run locally; the other providers are called over HTTP.
Schema migrations (embedded goose SQL scripts) are applied automatically on boot, so a fresh deployment is searchable with no extra step. To apply them explicitly instead:
make migrate-up # runs server/cmd/migrateConfiguration lives in server/config/config_<APP_ENVIRONMENT>.yaml (dev by
default); SQLITE_DSN can be overridden by the SQLITE_DSN environment variable. Each provider call is bounded by a context deadline — the shared
search_timeout_in_second, overridden per provider by
search_timeouts_in_second. There is no separate HTTP-client timeout: those
deadlines govern the whole upstream call, body read included. The daily
Project Gutenberg catalog download uses its own client with a 10-minute
budget. The books search knobs:
books:
default_genre: "non-fiction" # applied when the request omits genre
default_limit: 20 # applied when the request omits limit (max 50)
cache_ttl_in_second: 300 # identical successful searches come from memory for this long (0 disables)
middleware:
rate_limit:
enabled: true # per-client throttle on the API routes
requests_per_second: 5
burst: 20
providers:
search_timeout_in_second: 30 # default per-provider context timeout
search_timeouts_in_second: # per-provider overrides of that timeout
open_library: 20
doab: 15
gutendex: 25
internet_archive: 25
library_of_congress: 20
wikidata: 30
project_gutenberg: 10SQLite is opened in WAL mode with a 10s busy timeout so reads keep serving while the catalog refresh writes.
GET /metrics reports in-memory funnel counters — searches_total,
searches_cached_total, per-provider provider_status outcomes,
access_clicks, and events_rejected_total. They reset on restart and hold
no personal data. The frontend posts a small
{"type":"access_click","provider":...} beacon to POST /books/events when
a reader opens a book's read link, so the search→read click-through rate —
the launch success signal — is measurable end to end.
API routes (/books/*, /example) are rate limited per client address; the
probes and /metrics are exempt so monitoring cannot be locked out. Client
IPs are taken from the connection address (trusted proxies are disabled), so
a spoofed X-Forwarded-For cannot rotate limits. If you deploy behind a
reverse proxy, configure gin's trusted proxies accordingly.
The search cache stores the fused result pool per query and serves page
navigation from it for books.cache_ttl_in_second (default 300s), which
absorbs bursts without hammering rate-limited upstreams. Any search in which
at least one provider succeeded is cached; a fully failed search is never
cached. Cached responses report the provider statuses observed when the pool
was built, and the TTL is when the next request retries every provider.
- DOAB returns HTTP 403 from some deployment networks ("Your address is not allowed to access this API") — their API gate blocks client IPs. The provider uses DOAB's current DSpace 7 discover API; re-test from another network before assuming a code problem.
- loc.gov returns HTTP 403 (Cloudflare "Just a moment" challenge) for datacenter IPs; browser-like User-Agents do not help. May work from other networks.
- Gutendex occasionally times out from some networks; it has a 25s configured budget and recovers on retry.
- Deep pagination fetches the requested page plus a lookahead window and caches the fused pool, so page navigation within the pool is served from memory; pages beyond the pool (or a new query) re-fetch by design, which is the price of stateless, stable RRF ordering.
- Open Library sometimes serves blank/white cover images that load successfully, so the frontend's broken-image fallback cannot detect them.
Install dependencies and start Vite:
make web-install
make web-devThe frontend listens on http://localhost:5174 and proxies /api/* to the Go
server. The API client can be pointed at another server with
web/.env.local:
VITE_API_BASE_URL=/apiFrontend checks:
make web-lint
make web-typecheck
make web-test
make web-buildmake test # go test -short -race ./...
make lint # golangci-lint run ./...
make fmt # gofmt