agentlas
Marketplace
AgentRepair required

Runtime Evidence Evaluator

by Agentlas
Agent explainer

What this agent actually does

01

Job

같은 작업을 Claude, Codex, Gemini, API에 돌려 보고 어느 쪽 답이 더 나은지 객관적으로 가려야 할 때 씁니다. GitHub 이슈 링크나 제품 스펙 파일, 런타임 리포트를 주면 모델별 답을 나란히 놓고 품질을 채점해요. 채점한 근거는 트레이스 링크로 다 걸어 두고, 모델끼리 답이 너무 엇갈리면 그냥 넘어가지 않고 멈춰서 경고합니다. 끝나면 추적 가능한 JSON 증거 스코어카드랑 마크다운 증거 리포트를 레포 안에 남겨 줘요.

02

Tool use

If a run needs a plugin or external API, it asks for access first and uses it only within the approved scope.

03

Result

품질 점수와 트레이스 링크를 담은 eval-scorecard.json · 근거를 풀어 적은 eval-report.md와 trace-index.json

Best for

What it's good for

여러 모델 답을 근거까지 붙여 공정하게 비교하고 싶은 평가 엔지니어
어떤 런타임을 쓸지 데이터로 정하고 싶은 1인 빌더
모델 간 답 차이가 큰 케이스를 놓치지 않고 잡고 싶은 소규모 팀
What's inside

What's in this agent

1 skill1 agent1 memory2 mcp1 command
Outputs

What it produces

품질 점수와 트레이스 링크를 담은 eval-scorecard.json
근거를 풀어 적은 eval-report.md와 trace-index.json
멈췄을 때 이유를 적은 failure-summary.json, 과거 점수를 켜면 score-history.json
Prerequisites

Before you start

Claude, Codex, Gemini, API 실행 결과 (GitHub 이슈 링크나 스펙 파일, 런타임 리포트 파일)
GitHub만, 로컬만, 또는 섞인 형태로 둔 소스 위치
결과를 받을 위치 (레포 안 JSON과 마크다운 리포트)
Safety

What it can touch

Access
Files: scoped
Network: none
External API: yes
Be careful with
리포트를 레포 안에 새 파일로 쓰므로 덮어쓰기 전에 확인합니다
외부 도구나 클라우드 호출이 필요하면 먼저 물어봐요
소스가 없거나 모델 답이 너무 엇갈리면 추측 없이 멈춥니다
ONTOLOGY CHIPS

Operational experience and taste compatible with this agent

Hiring the agent and selecting an experience chip are separate decisions. Only verified exact-release matches appear, and none is purchased or attached automatically.

No publicly verified chip is available for this agent yet.
Sign in to create an attachment approval.

Viewing never purchases, attaches, or changes permissions.

Sign in
Safety

Inspect everything before it runs

A security scan runs before publish or install, and Agentlas never hosts or proxies models — it runs on your own account and keys.