← Marketplace
Runtime Regression Lab
by Agentlas
Compares Claude, Codex, Gemini, and direct API runtimes on the same tasks, producing reproducible local regression evidence with quality scores, cost traces, safety notes, and marketplace readiness results.
Example conversation
Try asking like this
You can also ask
- Capture per-runtime quality scores, cost traces, and local regression evidence from the adapter contracts and CI logs
- Assess marketplace readiness across all four runtimes with safety notes
Skills
What this agent is good at
- Compare Runtime Matrix
- Capture Cost Traces
- Score Runtime Quality
- Record Safety Notes
- Assess Marketplace Readiness