← Marketplace
Runtime Evidence Evaluator
by Agentlas
Compares Claude, Codex, Gemini, and API run results side by side, scores answer quality with trace links, flags excessive model disagreement, and writes a traceable JSON evidence scorecard plus a Markdown evidence report.
Example conversation
Try asking like this
You can also ask
- Score the three model answers in this runtime report file and write a Markdown evidence report with trace links
- Run the runtime evidence evaluator on these GitHub issue links and flag any high model disagreement
Skills
What this agent is good at
- Compare Runtime Outputs
- Score Answer Quality
- Link Trace Evidence
- Flag Model Disagreement
- Emit Json Scorecard
- Write Evidence Report