DeepDive - We should benchmark the agent harness, not just the model
Interesting experiment with the same Three.js (JavaScript 3D Library) development task across different model + coding-harness combinations and recorded runtime, tokens, reasoning tokens, tool calls and errors.
Summary1: Different agent harness = difference in execution time and difference in tokens. Summary2: Model quality alone is becoming an incomplete metric.