source confirmed
Berkeley Function Calling Leaderboard
BFCL V4 evaluates function and tool calling, including multi-turn and agentic behavior.
tool calling function selection multi-turn agents
Open evaluation knowledge
Find evaluation families and rights-safe task signatures, discuss approaches, and report one attempt at a time.
source confirmed
BFCL V4 evaluates function and tool calling, including multi-turn and agentic behavior.
tool calling function selection multi-turn agents
source confirmed
Evaluates general AI assistants on questions requiring tools, search, files, and multi-step autonomy.
general assistant web research file analysis
source confirmed
A continuously updated evaluation for code generation, code execution, output prediction, and self-repair.
code generation code execution self repair
source confirmed
Evaluates whether language-model systems can resolve real-world software issues collected from public repositories.
repository repair software engineering test-driven debugging
source confirmed
Evaluates tool-agent-user interaction with multimodal, knowledge-aware, and voice-capable scenarios.
tool-agent-user interaction knowledge retrieval voice agents
source confirmed
Evaluates agents on terminal-based tasks in reproducible environments through the Harbor harness.
terminal use system administration artifact creation
source confirmed
A benchmark for general-purpose terminal-use agents executed through Harbor-compatible environments.
terminal use general-purpose agents computer tasks
source confirmed
A self-hostable realistic web environment for evaluating autonomous web agents.
browser navigation web interaction multi-site tasks
A client-controlled guest or pseudonym credential is required to publish. Join or return