Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
-
Updated
Aug 21, 2026 - Python
Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
Score the trustworthiness of outputs from any LLM in real-time
🐀 Fuzz your verifier before an RL agent does. Static + dynamic LLM security auditor to detect reward-hacking in RL post-training environments (OpenEnv, verifiers-spec, Gymnasium).
Open Arena: SLM verifiers across observability platforms, dataset and model catalogs, and value scenarios for LLM, agentic and harness evals.
An RL Enviorment for AES Inversion
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
A verifiers RLM environment for testing whether adaptive recursive search outperforms brittle manual RAG choreography on long synthetic corpora.
Verifiable RL environments for corporate law & governance — deterministic reward, no LLM judge, fully synthetic worlds.
An open reinforcement-learning (RL) environment that trains LLM agents to use the current fact, not the stale one — verifiable reward for temporal fact-currency, built on verifiers / prime-rl (GRPO, LoRA).
Reproducible verifier audits, datasheets, agreement metrics, and release gates
RLVR coding environment for training and evaluating LLM agents on automated code-fixing tasks.
Deterministic SRT captioning environment for LLM evaluation and reinforcement learning.
A verifiable RL environment for TRP ion-channel ligand pharmacology, built on Prime Intellect's Verifiers
Research proposal for verifier-gated on-policy distillation with explicit evaluation and claim boundaries
Verifiers hello world repo
Minimal verifier environments for testing agents.
RL environments for generating code whose correctness is checked by formal verification; first release: C + ACSL + Frama-C.
A verifiers RL environment that trains models to propose novel, evidence-grounded, falsifiable hypotheses. Rewards novelty with accountability.
Three small demos on verifiable evaluation: measured retrieval accuracy, a validation loop that bounces bad work, and an end-to-end pipeline. No API key, no GPU.
To associate your repository with the verifiers topic, visit your repo's landing page and select "manage topics."