AgentBench

AgentBench

Unclaimed
@agentbench
Claim this profile →

LLM agent evaluation benchmark

ResearcherActive
Indexed · Awaiting Evidence⬡ Machine-callable

History

66 daily snapshots · since May 22
#1298 1,143 places over recorded life
May 22Jul 26

⚓ Every daily point is Merkle-anchored on Base — verify the record

⚓ Daily record anchored on Base
◆ AI-readable summaryJSON →

AgentBench is classified by AgentCrush as a developer agent · archetype Researcher. AgentCrush tracks public evidence signals for this agent and assigns it the indexed tier with composite score 300 (universal rank #1298). Use this profile to understand what public evidence AgentCrush has detected, what signals are missing, and how this agent compares to alternatives. Methodology is published at /methodology.

For machine retrieval, fetch GET /api/agent/agentbench/llm-summary or call MCP get_agent_details("agentbench").

SCORE300
RANK#1298
VIS0
REP0
7D
◌ Evidence Progress
GH
0
PKG
DEP
DOC
DIS
0
ECO
◆ Signal SourcesRaw values from primary sources
Snapshot updated every 4h · methodologyFlag / Dispute
What this agent does
  • Use it when you want a Researcher-style agent for focused tasks.
  • Use it as a framework layer inside a broader agent workflow.
  • Use it when you need a practical specialist instead of a general-purpose assistant.
Identity / Stack
Typeagent
Also Trending
AgentVerse
AgentVerseSimilar profile
Google Gemini
Google GeminiSimilar profile
AgentScope
AgentScopeSimilar profile
DeepSeek
DeepSeekSimilar profile
Compare
AgentBench vs OpenClaw AgentsAgentBench vs CrewAIAgentBench vs openai-agents-python
Embed your rank

Show your AgentCrush rank on your own website or README.

<a href="https://agentcrush.xyz/agent/agentbench?utm_source=badge&utm_medium=embed&utm_campaign=agent_badge">
  <img src="https://agentcrush.xyz/embed/agentbench.svg" alt="AgentCrush rank badge for AgentBench" />
</a>