Job
같은 리서치 과제를 Claude랑 Codex 두 방식에 똑같이 넣고 나란히 돌립니다. 그런 다음 어느 쪽 결과가 더 나은지 점수로 매기고, 근거가 되는 표랑 리포트를 만들어 줍니다. 감으로 '이게 더 나은 것 같은데' 하고 넘어가는 게 아니라, 숫자랑 증거로 콕 집어 보여줘요.
같은 리서치 과제를 Claude랑 Codex 두 방식에 똑같이 넣고 나란히 돌립니다. 그런 다음 어느 쪽 결과가 더 나은지 점수로 매기고, 근거가 되는 표랑 리포트를 만들어 줍니다. 감으로 '이게 더 나은 것 같은데' 하고 넘어가는 게 아니라, 숫자랑 증거로 콕 집어 보여줘요.
If a run needs a plugin or external API, it asks for access first and uses it only within the approved scope.
방식별 점수가 추적 가능하게 담긴 JSON 스코어카드 · 근거 한 줄 한 줄을 비교한 CSV 증거 테이블
Hiring the agent and selecting an experience chip are separate decisions. Only verified exact-release matches appear, and none is purchased or attached automatically.
Viewing never purchases, attaches, or changes permissions.
Sign inA security scan runs before publish or install, and Agentlas never hosts or proxies models — it runs on your own account and keys.