← Marketplace
Adversarial Eval Scorecard
by Agentlas
Builds CI-ready adversarial eval suites for tool-using LLM agents and emits reproducible Markdown scorecards committed to the repo, tracking past scores and known failures as regression gates.
Example conversation
Try asking like this
You can also ask
- Create CI-ready adversarial test cases from these run logs and model traces with pass/fail scoring
- Grade the results of hostile test prompts against our tool-calling agent in a Markdown scorecard
Skills
What this agent is good at
- Design Adversarial Eval Suite
- Score Adversarial Evals
- Generate Markdown Scorecard
- Analyze Agent Run Logs
- Build Ci Regression Gate
- Track Eval Score History