450 Musing
19 Sept 2026
Notes From Fucking Around With Vision, Text, and Audio Models
The bulk of my vibe-coding experiments over the last few months have involved audio and vision, so I’ve ended up trying an unreasonable number of multimodal models. This is a small guide based entirely on what I’ve actually used. If you have a project where you need vision capabilities for text extraction, models for…
keep reading — 14 min ↗