← Marketplace
Eval Memory Ledger
by Agentlas
Reviews new adversarial benchmark runs against past results, producing a traceable freshness report, a conflict ledger of contradicting findings, and a machine-readable evidence manifest.
Example conversation
Try asking like this
You can also ask
- Generate a freshness report and machine-readable evidence manifest for the new benchmark runs
- Match the new eval results against our memory ledger, then update the conflict ledger with contradictions
Skills
What this agent is good at
- Audit Benchmark Freshness
- Maintain Conflict Ledger
- Build Evidence Manifest
- Track Eval Run Provenance
- Deliver Run Summary