Jev on the WebMCP benchmark
Jev picks the tool, a fast LLM fills the arguments: 49 of 49 tasks solved.
Idan Levin’s team ran Jev on their open WebMCP benchmark. Jev picks each tool the website exposes, and Mercury 2.5 writes the arguments. That setup solved 49 of 49 tasks, at about 112 times lower model cost than GPT-6 Astra with computer use and code execution. Jev driving the page directly, through a modified Ultrafast, solved 25 of 49.
