
Video search and understanding require moving beyond simple metadata or transcript-based approaches to natively process the rich, spatial-temporal information inherent in video. Twelve Labs addresses this by utilizing specialized foundation models: Marengo, which creates vector embeddings for efficient search, and Pegasus, which converts raw video into structured, hyper-markup language. Unlike general-purpose multimodal LLMs that rely on sparse image snapshots, these models analyze video as a continuous stream, enabling precise retrieval and reasoning across massive, unlabeled archives. This technology transforms workflows in sectors like security, sports analytics, and documentary production by automating the extraction of qualitative insights and patterns. As enterprises seek more cost-effective, task-specific solutions, these specialized models provide a scalable alternative to general-purpose frontier models, allowing for deeper, context-aware analysis of visual data that was previously inaccessible or too labor-intensive to process.
Sign in to continue reading, translating and more.
Open full episode in Podwise