Pi agent is just better
Composio ran DeepSeek V4 Flash through 8 agent harnesses on 30 agentic tasks, and as per the report Pi agent came out on top in almost every metric. Here is the leaderboard for tasks passed.
Pi agent passed 20 of 30 tasks, 3 more than the next best harness (Oh My Pi at 17). Claude Code, Codex, and Deep Agents tied at 16, and the rest were at 14-15.
Cost per successful task is where Pi pulls ahead the most. It came in at $0.028, the cheapest of all 8 harnesses, while Claude Code at the other end cost $0.195, almost 7x more. Pi's median time per task was 132 seconds, only Claude Code and OpenCode were slightly faster.
The same model delivered 47-67% task success, cost $0.019-$0.104 per task, and took 122-272 seconds median time, depending on the harness. Their closing point is a good one. Benchmark the model-harness pair you will actually use, not the model in isolation.
Nice to see Pi win a third party eval like this, since I run Pi with DeepSeek as my default setup anyway.