Engineering / Evaluation Infrastructure

Private Grading for a Coding-Agent Benchmark

Grading infrastructure for JaseciBench: public CI checks, a private repository of hidden tests, and scores posted back to the pull request.

Problem If a benchmark publishes its tests, AI coding agents can be tuned to pass exactly those tests. JaseciBench keeps its grading tests private.

publicprivatesubmissionCI gateshiddentests+ oraclescore only
Schematic

Overview

JaseciBench evaluates language models and AI coding agents on Jac, an open-source, Python-like programming language that is almost absent from their training data. It measures three layers: single-shot generation, an agent repair loop, and end-to-end app delivery. This page covers how submissions are graded.

Architecture

 submission (fork + PR)            public repo                      private vault
 ─────────────────────        ────────────────────────       ──────────────────────────
  agent's changes      ──▶    ci.yml                          hidden tests
                              · jac type check     ──pass──▶  reference solutions
                              · baseline tests                scoring oracle
                                                                   │
  score table          ◀──    PR comment  ◀── grade-pr.yml ◀───────┘
                              (maintainer-triggered)

Key engineering decisions

  • Hidden tests in a separate repository. Hidden tests, reference solutions and the scoring oracle are in a private repository. The public repository has the tasks and baseline tests only.
  • Required checks first. Each task must pass jac check and the existing public tests before hidden tests count, so a submission cannot score points by breaking working code.
  • Maintainer-triggered grading. Official grading runs only when a maintainer starts the workflow, and the results are posted as a comment on the pull request.
  • Local grading. A scripts/grade command starts the same private workflow and prints the score table, so submitters get quick feedback without seeing the hidden tests.

Scope

The suite includes a Jac port of HumanEval (164 tasks) and a reference Jaseci application with three agent tasks, plus a Jac-native leaderboard.

Evidence

All engineering