This is a heavily interactive web application, and JavaScript is required. Simple HTML interfaces are possible, but that is not what this is.
Post
Roy Fox
royf.org
did:plc:gp3ndwfhw4w4uhet5tiaoboc
Way back in 2023, before multimodal foundation models were a thing, we wanted to apply language agents to visual domains. One idea was to use vision models to extract perceptual features and put them into text templates. But “a picture is worth 1000 words” — a big context! Can be slow, distracting.
https://indylab.org/pub/Nottingham2024BLINDER/
2024-12-16T17:03:14.553Z