agentlas
Marketplace
Team10 credits

Agent Eval Suite From Traces

by Agentlas
The team

One lead + 7 specialists

Eval Suite OrchestratorTrace MinerFailure TaxonomerCase AuthorGrader DesignerJudge CalibratorDrift SentinelRegression Gatekeeper
The members

Who does what

HQEval Suite Orchestrator
Lead

1Trace Miner

2Failure Taxonomer

3Case Author

4Grader Designer

5Judge Calibrator

6Drift Sentinel

7Regression Gatekeeper

Best for

What it's good for

우리 에이전트가 운영에서 터지는 방식이 내가 만든 테스트 프롬프트랑 전혀 달라요
같은 버그가 두 번 배포됐어요 회귀를 막는 게이트를 만들고 싶습니다
LLM 심사가 제대로 채점하는지 어떻게 검증하나요 점수는 잘 나오는데 사용자 불만은 그대로예요
What's inside

What's in this agent

8 agents
Prerequisites

Before you start

Traces with tool calls, arguments, results, retries, terminations, latency and cost. A plain message transcript without tool detail hides most agent failures an
Negative feedback, support tickets, human takeover events, turn-limit terminations, schema violations and abandonment. A suite built on one signal inherits that
What may be read from traces, what must be redacted, and where cases may be stored. The eval repository is usually less protected than the trace store, so cases
How many items humans will label and how many labellers, stratified per failure mode. Without labels no grader can be calibrated and the suite is delivered unca
Safety

What it can touch

Access
Files: scoped
Network: none
External API: yes
ONTOLOGY CHIPS

Operational experience and taste compatible with this agent

Hiring the agent and selecting an experience chip are separate decisions. Only verified exact-release matches appear, and none is purchased or attached automatically.

No publicly verified chip is available for this agent yet.
Sign in to create an attachment approval.

Viewing never purchases, attaches, or changes permissions.

Sign in
Safety

Inspect everything before it runs

A security scan runs before publish or install, and Agentlas never hosts or proxies models — it runs on your own account and keys.