Notes from messing around with vision, text, and audio models
The bulk of my vibe-coding experiments over the last few months have involved audio and video, or rather audio and vision, and I’ve been experimenting pretty heavily with multimodal models. So this is just a small guide based on what I’ve used. If you have some project where you need vision capabilities for text…
keep reading — 14 min ↗