Evaluation Infrastructure
Private Grading for a Coding-Agent Benchmark
If a benchmark publishes its tests, AI coding agents can be tuned to pass exactly those tests. JaseciBench keeps its grading tests private.
2026
Engineering
Case studies of systems I have built, maintained and fixed, grouped by system. Each one covers the problem, the design decisions and the public pull requests behind them.
Evaluation Infrastructure
If a benchmark publishes its tests, AI coding agents can be tuned to pass exactly those tests. JaseciBench keeps its grading tests private.
2026
Compilers / Languages
The compiler’s map of the paths a program can take was wrong for several common statements, so the errors and warnings built on it were wrong too.
Jaseci Labs · Mar 2026
AI Systems
AI coding assistants know little about Jac, a new language, and in my app-building runs smaller models struggled when shown every tool at once.
Jaseci Labs · Mar to Apr 2026
Developer Infrastructure
Deploys could run the wrong Jac version, fail minutes in after databases were already created, and undo a bad release only by rebuilding it.
Jaseci Labs · 2026