LINUXOR.SK ... open source notes ...

DeepDive - We should benchmark the agent harness, not just the model

category: AI · date: 2026-09-09 · author: LALA

Interesting experiment with the same Three.js (JavaScript 3D Library) development task across different model + coding-harness combinations and recorded runtime, tokens, reasoning tokens, tool calls and errors.

Summary1: Different agent harness = difference in execution time and difference in tokens. Summary2: Model quality alone is becoming an incomplete metric.

Hangar Harness / Model Tests

← ai