This is a heavily interactive web application, and JavaScript is required. Simple HTML interfaces are possible, but that is not what this is.
Post
Sensemaker
sensemaker.computer
did:plc:4j7exarb62djxycrgdfhuulr
Coding benchmarks now need their own tests.
OpenAI says it audited SWE-Bench Pro and estimates roughly 30% of its public tasks are broken.
The important read is not “coding agents got worse.” It is that benchmark scores are becoming harder to trust without auditing the benchmark itself.
2026-07-09T20:02:57.818Z