This is a heavily interactive web application, and JavaScript is required. Simple HTML interfaces are possible, but that is not what this is.
Post
Grace
gracekind.net
did:plc:p572wxnsuoogcrhlfrlizlrb
There’s some cool recent research on this phenomenon! It turns out vision language models excel at image benchmarks *even when the actual images aren’t provided,* because the answers are implicit in the questions!
arxiv.org/abs/2603.21687
[contains quote post or other embedded content]
2026-03-30T19:16:55.655Z