TwelveLabs Pegasus 1.6 — egocentric video understanding for physical AI
Pegasus 1.6 is TwelveLabs' first video-understanding model built for first-person (egocentric) footage, turning raw wearable or teleoperation video into timestamped action labels, dense captions and quality scores for robot training data.
Why it matters
Automates physical AI's biggest bottleneck — video annotation (the public Ego-Exo4D set took ~155 human hours per hour of video) — with five purpose-built workflows including action segmentation, dense captioning, quality scoring, search/curation and consent flagging, at $1.75 per hour of video.
What you could build with it
A factory-automation startup could mount action cameras on line workers and use Pegasus 1.6 to convert a week of footage into a searchable, timestamped SOP library, automatically flagging steps that deviate from the documented procedure.
Does it hold up?
Shipped as a live API with published pricing and SDK upgrades — immediately usable for teams sitting on footage; third-party usage evidence is still thin since the launch is five days old.
Built with TwelveLabs Pegasus 1.6
- The Robot Report: Pegasus 1.6 brings video understanding to physical AItherobotreport.com · Interview with TwelveLabs CEO Jae Lee: the model needs no proprietary cameras, and TwelveLabs pitches it as the video-understanding layer that must come before feeding footage into robot policies.
- SLOP TV News: TwelveLabs Pegasus 1.6 reads first-person video at $1.75 an hoursloptvnews.com · Pricing analysis: $1.75/hr of video, $3 per 1M image tokens, $7.50 per 1M output tokens — identical to Pegasus 1.5's rates — and a note that Pegasus 1.6 describes and organizes video rather than driving robots itself.
- DigitalToday: TwelveLabs launches Pegasus 1.6, expands beyond media to physical AIdigitaltoday.co.kr · Walk-through of the five workflows (action segmentation/labeling, dense captioning, quality scoring, search/curation, consent and compliance flagging) with a worked example of structuring a worker's pick-tool-finish sequence.
Learn more
First spotted on article: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71JetBrains Mellum 2.1: 12B MoE coding model trained for agentsJetBrains released Mellum 2.1, a 12B MoE model (2.5B active parameters, 128K context) under Apache 2.0, trained mainly…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.