Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
-
Updated
Sep 11, 2026 - Python
Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
Catch AI code hallucinations—fabricated packages, invented APIs, and contradicted behavior—deterministically, without trusting another model.
Turn Chaos Into Structure. A Type-Safe AI Agent that extracts valid JSON from unstructured data using PydanticAI, FastHTML, and Gemini 2.5.
Interactive Phoenix LiveView demonstrations of the Crucible Framework - showcasing ensemble voting, request hedging, statistical analysis, and more with mock LLMs
Code for the ICML 2026 Main Track Paper "Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement"
Four tiny local LLMs decide a pixel-art character's mood, pose and dialogue instead of a hand-written state machine. 72% of the arbiter's vetoes cited rules that don't exist, and every test still passed. A small study in valid outputs with invented justifications. Runs fully offline.
Three small LLMs, one CPU, no cloud. Rigorous benchmarking, structured output validation, and head-to-head quality scoring via Ollama + FastAPI.
A Python to Prolog pipeline that bounds Elasticsearch LLM errors by grounding ES diagnostics analysis in verified metrics.
Companion project for the TechnologyDig Academy tutorial on building reliable generative programs with Mellea.
Collection of LLM failure modes used on failmodes.com
Reference implementation of CAAF — three-pillar agent framework with monotonic convergence.
TypeScript eval harness for measuring whether Grok answers stay grounded in source evidence
Extracts a driver profile from a call transcript behind a validating schema gate, then screens a load board before ranking by effective rate per mile.
Map where your bolted-on AI feature breaks before customers do. A free Claude Code tool: fragility map, reliability score across six dimensions, ranked gaps, and a 30-day plan. Built by a threat-intel practitioner.
CrucibleFramework: A scientific platform for LLM reliability research on the BEAM
Official implementation of Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation.
Profile README. AI engineer working on agentic LLM systems, deterministic guardrails, and how models fail in long interactions.
When the agent breaks, which layer dropped the ball? Operator-facing failure-attribution taxonomy for AI agent estates: ten-way dictionary, MAST + AgentRx crosswalks, postmortem template. Part of the Spine catalog.
Public artifact bundle for the preprint 'Lightweight Evaluation and Operational Scorecards for Tool-Using AI Agents'
Make small local LLMs reliable at multi-step agent workflows. Constrained decoding + verify-before-commit: 96.5% end-to-end vs 2.0% unguarded on a 50-step run. Zero dependencies, runs on-prem.
To associate your repository with the llm-reliability topic, visit your repo's landing page and select "manage topics."