Engineering / Evaluation Infrastructure
Private Grading for a Coding-Agent Benchmark
Grading infrastructure for JaseciBench: public CI checks, a private repository of hidden tests, and scores posted back to the pull request.
Problem If a benchmark publishes its tests, AI coding agents can be tuned to pass exactly those tests. JaseciBench keeps its grading tests private.
Overview
JaseciBench evaluates language models and AI coding agents on Jac, an open-source, Python-like programming language that is almost absent from their training data. It measures three layers: single-shot generation, an agent repair loop, and end-to-end app delivery. This page covers how submissions are graded.
Architecture
submission (fork + PR) public repo private vault
───────────────────── ──────────────────────── ──────────────────────────
agent's changes ──▶ ci.yml hidden tests
· jac type check ──pass──▶ reference solutions
· baseline tests scoring oracle
│
score table ◀── PR comment ◀── grade-pr.yml ◀───────┘
(maintainer-triggered)
Key engineering decisions
- Hidden tests in a separate repository. Hidden tests, reference solutions and the scoring oracle are in a private repository. The public repository has the tasks and baseline tests only.
- Required checks first. Each task must pass
jac checkand the existing public tests before hidden tests count, so a submission cannot score points by breaking working code. - Maintainer-triggered grading. Official grading runs only when a maintainer starts the workflow, and the results are posted as a comment on the pull request.
- Local grading. A
scripts/gradecommand starts the same private workflow and prints the score table, so submitters get quick feedback without seeing the hidden tests.
Scope
The suite includes a Jac port of HumanEval (164 tasks) and a reference Jaseci application with three agent tasks, plus a Jac-native leaderboard.