Pi agent is just better

Composio ran DeepSeek V4 Flash through 8 agent harnesses on 30 agentic tasks, and as per the report Pi agent came out on top in almost every metric. Here is the leaderboard for tasks passed.

Tasks passed across all 8 agent harnesses on DeepSeek V4 Flash

Pi agent passed 20 of 30 tasks, 3 more than the next best harness (Oh My Pi at 17). Claude Code, Codex, and Deep Agents tied at 16, and the rest were at 14-15.

Cost per successful task is where Pi pulls ahead the most. It came in at $0.028, the cheapest of all 8 harnesses, while Claude Code at the other end cost $0.195, almost 7x more. Pi's median time per task was 132 seconds, only Claude Code and OpenCode were slightly faster.

Cost per successful task across all 8 agent harnesses

The same model delivered 47-67% task success, cost $0.019-$0.104 per task, and took 122-272 seconds median time, depending on the harness. Their closing point is a good one. Benchmark the model-harness pair you will actually use, not the model in isolation.

Nice to see Pi win a third party eval like this, since I run Pi with DeepSeek as my default setup anyway.