Skip to main content

Loading…

Skip to main content
HomeProjectsPostsApproachesStackResourcesContact
Justin Tsugranes LogoJustin Tsugranes Logo

Justin Tsugranes

HomeProjectsPostsApproachesStackResourcesContact

Stay in the loop

Occasional notes on what I'm building, lessons earned, and the studio behind it.

By subscribing, you agree to receive No spam. Unsubscribe in one click anytime. from Justin Tsugranes. No spam. Unsubscribe anytime. Privacy Policy

© 2026 Total Ventures LLC. All rights reserved.

Privacy PolicyTerms of ServiceCookie Policy
← All stack tools

Stack Tool · ai

Veo3

Veo3 is a text-to-video platform I'm tracking for future social content automation, specifically for short-form video on TikTok and Instagram Reels once core product monetization is established.

Veo3 offers a direct path to generating short-form video content from text prompts, a capability I've identified as critical for scaling social media presence without increasing manual video production overhead.

What it is

Veo3 is Google's text-to-video model, capable of generating high-quality, short video clips from natural language descriptions. It handles a range of styles, movements, and scene compositions, translating textual intent into visual motion. Unlike earlier generative video attempts, Veo3 demonstrates a stronger grasp of temporal consistency and object permanence within a clip, which is crucial for producing usable content. It's an API-first offering, which aligns with my agentic engineering approach, allowing programmatic control over video generation. This means I can feed it structured data or agent-generated scripts and receive a video artifact, rather than relying on a manual UI for each creation.

How I use it

Currently, I don't use Veo3 in production. Its integration is deferred. The plan is to deploy it for specific, high-volume social media channels where short, engaging video loops are essential. For Total Formula 1, this means generating quick race highlights or driver-focused clips from live commentary or data feeds. For Pregnancy Power Hour, it could produce educational snippets or motivational messages for TikTok and Instagram Reels, based on content pulled from Sanity. The workflow would involve an agent generating a script and prompt, sending it to Veo3, then taking the resulting video, potentially running it through Cloudinary for optimization and watermarking, before scheduling distribution. This entire process aims to automate the creation of dozens of unique video assets daily, without a human in the loop for initial generation. The focus remains on proving core product monetization before investing heavily in this content layer.

Why this over alternatives

I've evaluated several text-to-video platforms, including RunwayML and Pika Labs. Veo3 stands out primarily due to its API-first design and Google's underlying research capabilities. While RunwayML offers robust editing features and Pika Labs is strong for creative, stylized outputs, Veo3's strength lies in its potential for integration into a fully automated pipeline. My goal isn't to create cinematic masterpieces, but to generate high-volume, contextually relevant social content efficiently. The ability to integrate directly via API, without significant UI overhead or manual intervention, is non-negotiable for my studio's operating model. The quality-to-cost curve also appears more favorable for the specific use cases I have in mind, particularly for short, dynamic clips rather than longer narrative pieces. This allows me to scale content generation without scaling human video editors.

Where it falls short

Veo3, like all current text-to-video models, has limitations. The primary challenge is maintaining perfect temporal consistency over longer durations or with complex scene changes. While improved, artifacts and "jumps" can still occur, making it unsuitable for broadcast-quality content or anything requiring precise continuity. The cost model, while potentially favorable for my specific use case, is still a significant factor. Generating thousands of clips daily could become expensive quickly, which is why its deployment is tied to proven monetization. There's also a learning curve in prompt engineering to achieve desired outputs consistently. It's not a "magic button" — it requires iterative refinement of prompts and potentially post-processing steps. For example, generating specific facial expressions or precise brand-aligned visual elements for Inky might still require manual touch-ups or more advanced models.

FAQs

Is the output quality good enough for paid ads?
For short, dynamic social ads, yes. For high-production brand campaigns, likely not without significant manual post-production. It excels at quick, engaging loops.
How does the cost compare to hiring a junior video editor?
For high volume, short-form content, it will be significantly cheaper. For nuanced, creative work, a human editor is still more cost-effective. The break-even point depends on volume and desired fidelity.
Can it generate videos with specific brand elements or logos?
Not natively with perfect consistency yet. You'd likely need to overlay logos or brand elements in a post-processing step, potentially using Cloudinary transformations.

See how I run a multi-brand studio with this stack.

See the full stack →
Written by Justin Tsugranes, Founder, Total Ventures· Founder, Total Ventures · U.S. Army veteran (13 years) · M.M. Jazz Studies, University of South Carolina
Last reviewed May 9, 2026