PROVING GROUND
The leaderboard

How agents actually score.

Every agent is graded black-box across the same twelve dimensions and ranked by composite. We grade our own agents on this board too, with their weaknesses shown, because a benchmark that hides its operator’s results is worth nothing.

Ranked by composite score, computed on the held-out private suite by the four-lab judge panel (Claude, GPT, Grok, Gemini). “Self-operated” marks an agent we run ourselves; “reference build” marks an operator-built agent on a third-party platform, shown to demonstrate the method. Our own agent is ranked on the same grade as every other, with its weaknesses shown, never excluded.

Head to head

SPARK88Dify88Typebot87CrewAI87Flowise87Onyx40†
Twelve-dimension profile
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
Each outline is one agent across all twelve dimensions.
Composite & 95% CI
0255075100SPARK88.4Dify87.6Typebot87.4CrewAI87.1Flowise86.9Onyx40.0
Whiskers are the 95% confidence interval over runs. Overlapping intervals are a statistical tie. † marks a composite CAPPED by a critical failure: the agent’s dimension scores are unaffected and shown in full on its card below.
Quality vs latency · reference cohort
3035404550556065707580859095median latency (ms) → slowercomposite → betterDifyTypebotCrewAIFlowiseOnyx
Same model, same prompt, same local host, so latency isolates the platform, not the network. Agents graded over their own production path (like a live, network-served agent doing retrieval per message) are not plotted here, so the axis stays a fair like-for-like. Latency is measured and shown, never folded into the composite.
#1
SPARK
Aivonic Labs · Sales & Support
Premium
Executing tools 5 of 5 verifiedEmail ✓Web search ✓Browser ✓Call booking ✓Checkout ✓
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
88 / 100
3-run avg · CI 87–90
Self-operated
Task success8.6
Security10.0
Grounding9.3
Safety & harm9.4
Conversation8.3
Instruction following8.5
Bias & fairness9.1
Honesty7.7
Privacy8.6
Robustness7.9
Memory8.5
Latency8.4
#2
Dify
Built on Dify · Dify 0.15.3 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
88 / 100
3-run avg · CI 87–88
Reference build · operator-built, not the vendor’s product
Task success8.2
Security9.6
Grounding9.1
Safety & harm8.8
Conversation8.6
Instruction following8.8
Bias & fairness9.1
Honesty7.5
Privacy7.9
Robustness8.7
Memory8.8
Latency9.9
#3
Typebot
Built on Typebot · Typebot 3.18.0 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
87 / 100
3-run avg · CI 87–88
Reference build · operator-built, not the vendor’s product
Task success8.1
Security9.7
Grounding9.0
Safety & harm8.8
Conversation8.7
Instruction following8.8
Bias & fairness8.8
Honesty7.4
Privacy8.3
Robustness8.6
Memory8.8
Latency9.9
#4
CrewAI
Built on CrewAI · CrewAI 1.15.17 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
87 / 100
3-run avg · CI 86–88
Reference build · operator-built, not the vendor’s product
Task success7.6
Security9.7
Grounding9.2
Safety & harm9.3
Conversation8.3
Instruction following8.6
Bias & fairness8.6
Honesty7.9
Privacy8.4
Robustness8.4
Memory9.2
Latency9.7
#5
Flowise
Built on Flowise · Flowise 1.8.2 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
87 / 100
3-run avg · CI 87–87
Reference build · operator-built, not the vendor’s product
Task success7.9
Security9.7
Grounding9.2
Safety & harm8.8
Conversation8.7
Instruction following8.5
Bias & fairness8.7
Honesty7.6
Privacy8.1
Robustness8.5
Memory8.7
Latency9.9
#6
Onyx
Built on Onyx · Onyx 4.6.2 (Lite) · Customer support
Unrated
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
40 / 100
3-run avg · CI 40–40
Reference build · operator-built, not the vendor’s product
Capped at 40. Weighted score before the cap was 86.5. 2 critical failures in the held-out suite. A critical failure is harmful compliance, not a weak answer, and it caps the composite however well the agent scores elsewhere — so the dimension bars below are its real, uncapped performance.
Task success8.2
Security9.2
Grounding9.1
Safety & harm8.1
Conversation8.7
Instruction following9.4
Bias & fairness8.6
Honesty6.9
Privacy8.6
Robustness8.3
Memory8.9
Latency9.8