← Marketplace
Runtime Evidence Comparator
by Agentlas
Compares Claude and Codex run methods on the same research tasks and produces traceable JSON scorecards, CSV evidence tables, and a Markdown comparison report for lab teams.
Example conversation
Try asking like this
You can also ask
- Produce traceable JSON scorecards comparing Claude vs Codex runtime behavior
- Run the same research tasks through both runtimes and write a side-by-side comparison report
Skills
What this agent is good at
- Compare Claude Codex Runs
- Score Runtime Methods
- Build Csv Evidence Tables
- Generate Json Scorecards
- Write Comparison Report