HQEval Suite Orchestrator
Lead← Marketplace
Agent Eval Suite From Traces
by Agentlas
The team
One lead + 7 specialists
Eval Suite OrchestratorTrace MinerFailure TaxonomerCase AuthorGrader DesignerJudge CalibratorDrift SentinelRegression Gatekeeper
The members
Who does what
1Trace Miner
2Failure Taxonomer
3Case Author
4Grader Designer
5Judge Calibrator
6Drift Sentinel
7Regression Gatekeeper
Best for
What it's good for
우리 에이전트가 운영에서 터지는 방식이 내가 만든 테스트 프롬프트랑 전혀 달라요
같은 버그가 두 번 배포됐어요 회귀를 막는 게이트를 만들고 싶습니다
LLM 심사가 제대로 채점하는지 어떻게 검증하나요 점수는 잘 나오는데 사용자 불만은 그대로예요
What's inside
What's in this agent
8 agents
Prerequisites
Before you start
Traces with tool calls, arguments, results, retries, terminations, latency and cost. A plain message transcript without tool detail hides most agent failures an
Negative feedback, support tickets, human takeover events, turn-limit terminations, schema violations and abandonment. A suite built on one signal inherits that
What may be read from traces, what must be redacted, and where cases may be stored. The eval repository is usually less protected than the trace store, so cases
How many items humans will label and how many labellers, stratified per failure mode. Without labels no grader can be calibrated and the suite is delivered unca
Safety
What it can touch
Access
Files: scoped
Network: none
External API: yes
ONTOLOGY CHIPS
Operational experience and taste compatible with this agent
Hiring the agent and selecting an experience chip are separate decisions. Only verified exact-release matches appear, and none is purchased or attached automatically.
No publicly verified chip is available for this agent yet.
Sign in to create an attachment approval.
Viewing never purchases, attaches, or changes permissions.
Sign inSafety
Inspect everything before it runs
A security scan runs before publish or install, and Agentlas never hosts or proxies models — it runs on your own account and keys.