Editorial workflow yang menghubungkan video talking head, storyboard, start frame, end frame, dan motion B-roll
Tutorial7 menit bacaIntermediate

Cara Bikin B-roll Talking Head dengan Gemini, Image Generation, dan Google Flow

Workflow lengkap mengubah transkrip talking head menjadi storyboard visual 9:16, still image yang konsisten, lalu motion B-roll di Google Flow.

Sumber ↗ · Tutorial based on Mustafa's own talking-head B-roll workflow. Header and example media approved by Mustafa for publication.

B-roll talking head tidak harus dimulai dari video AI. Workflow ini dimulai dari video utama, transcript dengan timestamp presisi milidetik dari Gemini, still image 9:16 untuk setiap visual beat, lalu image-to-video motion di Google Flow.

Contoh ini memakai video tentang Hermes, OpenClaw, Obsidian, VPS, Mac mini, dan komputer. Still image dibuat sebagai end frame yang siap dianimasikan, lalu clip motion dipasang kembali sebagai B-roll di timeline talking head.

  1. 1. Mulai dari video talking head

    Rekam video utama lebih dulu. Audio dan timing video menjadi source of truth untuk seluruh visual.

  2. 2. Minta Gemini membuat transcript dan timestamp

    Upload video utama ke Gemini. Timestamp ini menjadi perkiraan durasi tiap B-roll supaya generation tidak kepanjangan atau kependekan dan credit Google tetap hemat. Durasi final tetap bisa dipotong manual di editor.

    Prompt Gemini untuk transkrip · text
    Transkrip video ini verbatim dalam Bahasa Indonesia. Pecah berdasarkan beat kalimat dan sertakan timestamp presisi milidetik dengan format [HH:MM:SS.mmm - HH:MM:SS.mmm]. Sertakan durasi tiap beat, speaker jika lebih dari satu, dan tandai bagian yang timestamp-nya hanya perkiraan. Jangan mengklaim presisi microsecond atau frame accuracy jika source video/provider tidak menyediakannya.

    Screenshot Gemini dengan video talking head dan prompt transkripsi bertimestamp
    Step 2 — video talking head di-upload ke Gemini, lalu ditranskrip dengan timestamp presisi.
  3. 3. Ubah transkrip menjadi visual beats

    Satu beat memuat satu ide visual. Pertahankan timestamp asli sebagai perkiraan durasi agar generation tidak boros credit; durasi final tetap dipotong manual di editor.

    Prompt storyboard visual · text
    Pecah transcript berikut menjadi visual beats untuk B-roll talking head 9:16. Tempel hasil transcript Step 2 pada bagian [TRANSCRIPT DI SINI]. Untuk setiap beat, tulis: timecode asli, durasi perkiraan, inti kalimat, satu metafora visual, objek wajib, arah gerak, area aman subtitle 20–25 persen bagian bawah, start frame, end frame, dan risiko continuity. Jangan membuat visual dashboard generik dan jangan mengubah timecode. Durasi hanya acuan hemat credit; durasi final boleh dipotong manual di editor.
    
    [TRANSCRIPT DI SINI]

Kunci workflow: Gemini menangani transkrip dan timing; image generator menangani komposisi still image; Google Flow menangani gerak antar-keyframe; editor menangani trim, placement, dan hubungan B-roll dengan talking head.

Jangan meminta image generator membuat frame abstrak lalu berharap Google Flow memperbaiki semuanya. Berikan end frame jelas, objek konsisten, arah gerak eksplisit, kamera locked, dan negative prompt untuk mencegah drift.

Sebelum publish, jalankan visual QA: aspect ratio, subtitle safe area, continuity, object readability, unwanted text, logo accuracy, dan artifact check.

  1. 4. Buat start frame Visual Part 1

    Start frame Visual 1 memakai blank canvas lama sebagai frame 0. Jangan ekstrak frame pertama dari video hasil render.

    Prompt image: start frame Visual 1 · text
    Create the START FRAME for Visual Part 1 as a completely blank 9:16 portrait canvas. Use full-frame warm off-white paper color #F8F6EF with extremely subtle fine paper grain only. No objects, no circles, no badges, no logos, no notebook, no arrows, no shapes, no text, no border, no shadows, no vignette, no gradients, no watermark. Flat even lighting. This is the clean first keyframe for compositing and animating real logo assets later.

    Blank warm off-white paper canvas sebagai start frame Visual 1
    Start frame Visual 1 — blank canvas yang dipakai sebagai frame awal motion.
  2. 5. Buat end frame Visual Part 1

    End frame memakai komposisi lama Visual 1. Ambil logo resmi dari `https://obsidian.md/images/obsidian-logo-gradient.svg`, `https://raw.githubusercontent.com/NousResearch/hermes-agent/main/website/static/img/logo.png`, dan asset OpenClaw dari repository resminya. Download dengan `curl -L --fail -o logo.svg <URL>`, rasterize SVG memakai `sips -s format png logo.svg --out logo.png` bila perlu, lalu composite ke badge kosong memakai layer image editor/FFmpeg. Jangan minta generator menggambar ulang logo: gunakan file logo asli sebagai layer deterministic, lalu export end frame PNG. Logo tidak digambar ulang oleh generator.

    Prompt image: end frame Visual 1 · text
    Create the END FRAME for Visual Part 1, the opening beat `00:00.000–00:03.950`, as a 9:16 vertical editorial paper-collage composition. Use warm off-white paper #F8F6EF, subtle paper grain, flat torn-paper layers, cobalt blue, tomato red, mustard yellow, and dark purple accents, imperfect black ink outlines, and tactile paper shadows. Show two separate circular paper badges: Hermes Agent in the upper-left and OpenClaw in the lower-right, connected through a small dark-purple faceted notebook in the center with hand-drawn black ink arrows. Keep lower 25 percent empty for subtitles and preserve generous safe margins. Use the real Hermes and OpenClaw logo assets supplied separately; do not redraw, stylize, morph, replace, distort, recolor, or hallucinate either logo. No invented words, fake UI, watermark, humans, neon, photorealism, or 3D CGI.

    End frame Visual 1 dengan logo Hermes dan OpenClaw pada badge yang terhubung ke notebook ungu
    End frame Visual 1 — Hermes dan OpenClaw terhubung melalui notebook pusat.
  3. 6. Animate start frame ke end frame di Google Flow

    Masukkan start frame dulu, lalu end frame, ke Google Flow. Set model `Omni`, aspect ratio `9:16`, output `720p`, dan gunakan durasi perkiraan dari beat agar credit tidak boros. Setelah generate, pilih `upscale` ke `1080p` saat download. Durasi final tetap bisa dipotong manual di editor.

    Prompt Google Flow: Visual 1 start → end · text
    Create an approximately 5-second 9:16 vertical editorial paper-collage animation using exactly two keyframes. Upload supplied Keyframe 1 first at frame 0: the completely blank warm off-white paper canvas. Upload supplied Keyframe 2 second at the final frame: the finished Visual 1 end frame. In Google Flow use model Omni, aspect ratio 9:16, output 720p. Animate from blank canvas into the exact supplied end frame: hold blank 0.0–0.7s; bring in the Hermes paper badge 0.7–1.5s; bring in the OpenClaw paper badge 1.4–2.2s; slide the purple notebook into center 2.0–2.8s; draw black ink lines from Hermes to notebook then notebook to OpenClaw 2.7–3.8s; settle all elements 3.8–4.5s; hold exact end frame 4.5–5.0s. Preserve the supplied real logos exactly. No logo redraw, morph, replacement, distortion, recolor, invented text, fake UI, watermark, camera movement, crop, neon, or 3D CGI. Keep lower 25 percent clean for subtitles. No music or voiceover; use subtle paper slide, paper tap, and ink-drawing sounds. After generation, download with `upscale` to 1080p if available.

    Screenshot Google Flow dengan dua keyframe dan prompt animasi video vertikal
    Google Flow — masukkan start frame dan end frame, lalu jalankan prompt motion.
  4. 7. Lihat hasil video generated Visual Part 1

    Hasil render motion Visual 1 berdurasi sekitar 4 detik. Pakai clip ini sebagai contoh output sebelum masuk tahap trim dan assembly.

    Keterangan video: Hasil video generated Visual Part 1 — clip 9:16 berdurasi sekitar 4 detik.

    Hasil video generated Visual Part 1 — clip 9:16 berdurasi sekitar 4 detik.
  5. 8. Trim dan pasang sebagai B-roll

    Edit manual di CapCut. Trim hasil render ke timecode beat `00:00.000–00:03.950`, letakkan di atas talking head, lalu cek subtitle safe area, logo readability, continuity, dan artifact.

    Screenshot CapCut dengan video talking head, subtitle, audio track, dan B-roll di timeline
    CapCut — pasang B-roll di atas talking head, sinkronkan dengan timecode, lalu cek subtitle dan audio track.
  6. 9. Ulangi workflow untuk semua visual tersisa

    Visual 1 menjadi template kerja. Untuk setiap beat berikutnya, pakai timecode sebagai perkiraan durasi supaya credit Google hemat. Pilih sendiri apakah beat perlu continuity dari visual sebelumnya atau menjadi scene independent. Buat still image sebagai end frame, buat atau pilih start frame yang sesuai, lalu render, trim manual, dan QA. Jangan copy komposisi Visual 1 mentah-mentah; copy struktur proses dan style saja.

    Prompt repeat workflow: semua visual tersisa · text
    Repeat the Visual 1 workflow for each remaining visual beat. Use the beat timecode and approximate duration from the storyboard as a credit-saving guide; trim final duration manually in the editor. For each beat, define one clear visual idea, required objects, motion direction, start frame, and end frame. Generate the still image as the end frame first. Use the previous beat’s approved end frame as the next start frame whenever continuity is intended; otherwise create a new blank or establishing start frame and state why. If a real brand logo appears, download the official asset from its source URL, rasterize SVG with `sips` when needed, composite the original logo deterministically into the end frame, and never ask the image or video generator to redraw it. Send exactly two keyframes to Google Flow, set the end keyframe at the final frame, render longer than the target if a hold is needed, then trim to the original timecode. Verify 9:16 framing, subtitle-safe lower 20–25 percent, object continuity, motion direction, logo fidelity, no invented text, no warping, and clean final hold. If QA fails, mark REVISE and do not use the clip.

  7. 10. Lihat hasil final setelah semua B-roll digabung

    Setelah semua motion B-roll selesai di-trim, gabungkan seluruh clip dengan video talking head, subtitle, dan audio di editor. Video Instagram ini adalah hasil final gabungan, bukan hasil motion Visual Part 1 saja.

    Hasil final video — talking head, seluruh B-roll, subtitle, dan audio sudah digabung lalu di-upload ke Instagram. Buka di Instagram