13 stages. One video. Zero guesswork.
Every waaave video runs through a fully automated pipeline — from niche research to posting. You step in at Stage 11 to review and approve. Everything else is handled.
Research
Stage 1 — Niche & Product
Your niche configuration and product details are read from onboarding. The pipeline validates that a product record exists, checks your credit balance, and creates a job record before any external API fires.
Stage 2 — Trend Research
Virlo Orbit scans within-niche creators for velocity-ranked formats. Virlo Satellite goes deep on the top outlier creators. A cross-niche analysis (Claude Sonnet 4.6 via Vertex AI) identifies 1–2 borrowed formats from adjacent categories that have structural mechanics applicable to your product.
Script
Stage 3 — Structural Decomposition
Claude deconstructs the winning format — not just the surface aesthetics but the underlying hook structure, pacing, information architecture, and CTA placement. The output is a layout JSON that becomes the blueprint for your video.
Stage 4 — Script Generation
Claude writes a hook-first script mapped to your product's specific benefit. Per-frame text overlay copy is included. The script includes a soft CTA that aligns with TikTok's native call-to-action patterns for your category.
Production
Stage 5 — Voice Synthesis
Your voice preference determines the synthesis path: Seed Audio 1.0 for English ambient, ElevenLabs multilingual for non-English, or ElevenLabs with your saved voice clone ID. Audio is uploaded to Cloud Storage — Convex stores the signed URL only.
Stage 6 — Visual Generation
One of six tracks is selected based on your format and available assets. Track A: LTX-Video 2.3 (default). Track B: Runway gen4_turbo (product footage). Track C: HeyGen avatar (avatar archetype). Track D: LTX-2.3 Fast (ambient b-roll). Track E: Reve 2.0 slideshow with Satori text overlays. Track F: Pinterest-sourced → Reve 2.0 reproduction (IP-safe). NSFW detection runs on every frame.
Stage 7 — Sound Matching
Virlo's sound velocity API returns the top trending audio for your niche. The sound ID is stored in your job record — it's applied natively at the TikTok posting stage, not embedded in the MP4 file. This is the correct approach per TikTok's API.
Stage 8 — Subtitles
AssemblyAI Universal-2 transcribes your voiceover with word-level timestamps. A custom SRT generator produces subtitle segments of max 7 words or 3 seconds. The SRT is uploaded to Cloud Storage for the FFmpeg compositing stage.
Stage 9 — Thumbnail
Reve 2.0 generates a product-focused background visual. If Satori is configured, hook copy is rendered as a PNG text overlay using your brand typography. The final thumbnail is composited in Stage 10.
Stage 10 — Compositing
A self-hosted FFmpeg worker composites all assets into a TikTok-spec MP4 (H.264, AAC, 9:16, 1080×1920). Subtitle burning, audio sync, thumbnail compositing, and optional watermarking (free-credit jobs) are all handled here.
Review
Stage 11 — Human Review← you step in here
The pipeline pauses at a Trigger.dev waitpoint. You review the video, see the Virlo trend data that generated it, and can edit hook + CTA copy inline without re-rendering. Three actions: Approve & Schedule, Regenerate (with format options), or Edit copy only.
Posting
Stage 12 — Scheduling & Posting
On approval, waaave posts the MP4 to TikTok via Zernio's certified integration using the official TikTok Content Posting API. The trending audio is applied natively at this stage. Your scheduled post time is respected. A post record is created in Convex.
Analytics
Stage 13 — Analytics Ingestion
A daily Convex cron polls Zernio's analytics API for every published post. Views, likes, shares, comments, engagement rate, and follower delta are written to your analytics records — each attributed to the Virlo trend that generated them.