Skip to content

Pinned Loading

  1. terminal-bench-1 terminal-bench-1 Public

    A benchmark for LLMs on complicated tasks in the terminal

    Python 2.6k 567

  2. harbor harbor Public

    Framework for evaluating and improving agents

    Python 4.7k 1.7k

  3. terminal-bench-2 terminal-bench-2 Public

    Shell 387 112

  4. terminal-bench terminal-bench Public

    Measuring and evolving with the frontier of agent work

    Python 545 413

  5. terminal-bench-science terminal-bench-science Public

    Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains

    Python 281 281

  6. awesome-harbor awesome-harbor Public

    A curated list of awesome Harbor ecosystem projects

    52 3

Repositories

Showing 10 of 19 repositories

Top languages

Loading…

Most used topics

Loading…