Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
-
Updated
Aug 18, 2026 - Shell
Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
VCR cassettes for agent trajectories: record agent runs as DAGs, replay the canonical path, only call the LLM for net-new paths
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Independent research on trajectory-aware AI-agent evaluation, delegated authority, control integrity, and failure-preserving reproducibility.
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Capture coding-agent sessions (Claude Code / Codex) at the source — trajectory + verifiable git environment — and score each session's training value (grounded × rich × focused). Local, deterministic, no model at runtime.
Add a description, image, and links to the agent-trajectories topic page so that developers can more easily learn about it.
To associate your repository with the agent-trajectories topic, visit your repo's landing page and select "manage topics."