Skip to content
Chandra179Public

About

all about books

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Trafae

Trafae is a book discovery service: it searches multiple open book catalogs at once, deduplicates the results, and fuses them into a single ranked list. It contains a Go API server and a React frontend.

App preview

Trafae searching "psychology" across open book catalogs: search, genre chips, provider status and fused results

Book search

GET /books/search fans a query out to every configured provider in parallel and merges the responses with Reciprocal Rank Fusion. Duplicate books across providers are collapsed into one result that keeps the per-source metrics. Every parameter is optional: search by topic (topic, repeatable), browse by genre alone (the genre doubles as the search term across providers when no topic is given), or combine both with the filters below.

Paging is stateless: pass page (1-based, default 1) alongside limit (page size). A page-N request fetches the requested page plus a few pages of lookahead from every provider, fuses and sorts that pool, slices it, and caches the pool for books.cache_ttl_in_second — page navigation within the cached pool is served from memory in milliseconds with zero upstream calls, and ordering stays stable across pages. Requests deeper than the cached pool re-fetch with fresh lookahead. Providers cap their fetch depth (Open Library, DOAB and loc.gov at 100 results; Internet Archive, Gutendex and Project Gutenberg at 200), so very deep pages run out of pool and has_more honestly turns false — with the frontend's limit of 24 that is around page 10.

Query param Meaning
topic search term, repeatable (max 8); optional
genre genre filter; defaults to default_genre, works standalone
page 1-based page number (max 20)
limit page size (default 20, max 50)
provider restrict to specific providers, repeatable
language two-letter language filter
min_year / max_year publication year range
min_popularity provider-specific popularity (e.g. downloads)
min_rating 0–5 rating floor

Supported providers:

Provider Popularity Ratings Notes
Open Library — 0–5 Rate limited; set open_library_contact_email for higher limits
DOAB — — Scholarly open-access books via the DSpace 7 API; links go to the DOAB record
Gutendex Downloads — Project Gutenberg via Gutendex
Internet Archive Downloads 0–5
Library of Congress — — Digital collections
Wikidata — — Community-supplied metadata via the Wikidata Action API
Project Gutenberg Downloads — Official catalog synced daily into SQLite

GET /books/providers lists each provider with the filters it supports, so clients can adapt queries to what will actually be served.

Examples:

curl "http://localhost:8081/books/search?topic=biology&genre=non-fiction&min_rating=4&limit=10"
curl "http://localhost:8081/books/search?genre=history&page=2&limit=24"

Backend

make run

The API listens on http://localhost:8081. Health, readiness, and funnel counters are available at /health, /ready, and /metrics.

On startup the server syncs the official Project Gutenberg catalog (a gzipped CSV) into SQLite and refreshes it every 24 hours. Searches against it run locally; the other providers are called over HTTP.

Schema migrations (embedded goose SQL scripts) are applied automatically on boot, so a fresh deployment is searchable with no extra step. To apply them explicitly instead:

make migrate-up   # runs server/cmd/migrate

Configuration

Configuration lives in server/config/config_<APP_ENVIRONMENT>.yaml (dev by default); SQLITE_DSN can be overridden by the SQLITE_DSN environment variable. Each provider call is bounded by a context deadline — the shared search_timeout_in_second, overridden per provider by search_timeouts_in_second. There is no separate HTTP-client timeout: those deadlines govern the whole upstream call, body read included. The daily Project Gutenberg catalog download uses its own client with a 10-minute budget. The books search knobs:

books:
  default_genre: "non-fiction"   # applied when the request omits genre
  default_limit: 20              # applied when the request omits limit (max 50)
  cache_ttl_in_second: 300       # identical successful searches come from memory for this long (0 disables)

middleware:
  rate_limit:
    enabled: true                # per-client throttle on the API routes
    requests_per_second: 5
    burst: 20

providers:
  search_timeout_in_second: 30   # default per-provider context timeout
  search_timeouts_in_second:     # per-provider overrides of that timeout
    open_library: 20
    doab: 15
    gutendex: 25
    internet_archive: 25
    library_of_congress: 20
    wikidata: 30
    project_gutenberg: 10

SQLite is opened in WAL mode with a 10s busy timeout so reads keep serving while the catalog refresh writes.

Launch funnel & protection

GET /metrics reports in-memory funnel counters — searches_total, searches_cached_total, per-provider provider_status outcomes, access_clicks, and events_rejected_total. They reset on restart and hold no personal data. The frontend posts a small {"type":"access_click","provider":...} beacon to POST /books/events when a reader opens a book's read link, so the search→read click-through rate — the launch success signal — is measurable end to end.

API routes (/books/*, /example) are rate limited per client address; the probes and /metrics are exempt so monitoring cannot be locked out. Client IPs are taken from the connection address (trusted proxies are disabled), so a spoofed X-Forwarded-For cannot rotate limits. If you deploy behind a reverse proxy, configure gin's trusted proxies accordingly.

The search cache stores the fused result pool per query and serves page navigation from it for books.cache_ttl_in_second (default 300s), which absorbs bursts without hammering rate-limited upstreams. Any search in which at least one provider succeeded is cached; a fully failed search is never cached. Cached responses report the provider statuses observed when the pool was built, and the TTL is when the next request retries every provider.

Known provider limitations

  • DOAB returns HTTP 403 from some deployment networks ("Your address is not allowed to access this API") — their API gate blocks client IPs. The provider uses DOAB's current DSpace 7 discover API; re-test from another network before assuming a code problem.
  • loc.gov returns HTTP 403 (Cloudflare "Just a moment" challenge) for datacenter IPs; browser-like User-Agents do not help. May work from other networks.
  • Gutendex occasionally times out from some networks; it has a 25s configured budget and recovers on retry.
  • Deep pagination fetches the requested page plus a lookahead window and caches the fused pool, so page navigation within the pool is served from memory; pages beyond the pool (or a new query) re-fetch by design, which is the price of stateless, stable RRF ordering.
  • Open Library sometimes serves blank/white cover images that load successfully, so the frontend's broken-image fallback cannot detect them.

Frontend

Install dependencies and start Vite:

make web-install
make web-dev

The frontend listens on http://localhost:5174 and proxies /api/* to the Go server. The API client can be pointed at another server with web/.env.local:

VITE_API_BASE_URL=/api

Frontend checks:

make web-lint
make web-typecheck
make web-test
make web-build

Checks

make test   # go test -short -race ./...
make lint   # golangci-lint run ./...
make fmt    # gofmt

About

all about books

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

Generated from Chandra179/lux