cap-cue

captions srt vtt subtitles timing
by @iberry420video
Produce a clean, timed caption track (SRT or WebVTT) from a narration script or corrected ASR: speech-aligned cues, 1-2 lines, ~32-42 chars/line, reading speed ≤~17 chars/sec, clean vs verbatim. Use when the user wants captions, subtitles, SRT, VTT, out-of-sync captions, or captions that flash too fast. Not burn-in/styling (hand to motion-craft), not kinetic word-by-word, not writing the spoken script (beat-script/hook-3s).
SKILL PACKAGE

Directory layout for Grok — SKILL.md plus scripts, references, and assets. 6 file(s).

Select a file

Edit in place, then Save (full package security scan). Use fullscreen for a larger workspace.

SYNCED SUMMARY (from SKILL.md)
# Cap Cue

If they cannot read it in time, it is not a caption. Cue the speech. Keep the line short. Then get out of the way.

Inspired by SkillMedev/skills captions-from-transcript (MIT). Original Grok remake. Complements beat-script (the spoken YouTube script), hook-3s (muted-proof first three seconds of a short), thumb-pack (the click image), shot-to-cut (assemble/render), motion-craft (burn-in, motion type, safe-area styling). Video captions only.

Read `references/cue-rails.md` before you time a line. Fill the three templates under `templates/`.

## When to use

- Need an SRT or WebVTT track from a narration script, a corrected ASR dump, or a talk-track with rough times
- Captions lag, lead, overlap, or flash faster than a reader can keep up
- Want speech-aligned cues: 1-2 lines, ~32-42 characters per line, reading speed at or under ~17 characters per second
- Choosing clean (drop fillers, keep meaning) vs verbatim (every um, for a record)
- Multi-speaker labels when the voices actually change

## When NOT to use

- Burn-in, kinetic word-by-word, lower-thirds, safe-area styling, brand motion type → use motion-craft
- Writing the spoken script, YouTube first-30s, retention beats → use beat-script
- TikTok / Reels / Shorts opener as spoken + frame-one + overlay → use hook-3s
- Thumbnail concepts → use thumb-pack
- Shot list, concat, render QA → use shot-to-cut
- Social-post captions (Instagram body copy, tweet threads) — different job, out of scope
- Phishing, malware pitches, fake medical cures, guaranteed-return finance, CSAM, sexual content involving minors

## Inputs

Gather what you can. Infer missing pieces and LABEL guesses. Ask at most two questions if critical data is missing.

- Script, corrected transcript, or ASR with any timestamps already present
- Video duration if known, and the destination (YouTube, Shorts, course player, web `<track>`)
- Format pick: SRT (editors / most uploads) or WebVTT (HTML5 track, web player)
- Mode: clean (default for video) or verbatim (they asked for a record)
- Speaker names if they supplied them. Do not invent identities.
- Language / on-screen UI terms that must stay exact

## Procedure

1. Read `references/cue-rails.md`. Refuse phishing, malware, fake cures, guaranteed-return finance, CSAM, and sexual content involving minors. Do not write the spoken script. Do not burn-in or animate type.
2. Fill `templates/source-card.md`. Lock mode (clean vs verbatim), format (SRT vs VTT), duration, and what time evidence you actually have. Label guessed times `[TIME: guess]`.
3. Segment speech into cues at clause boundaries. 1-2 lines, never 3. ~32-42 characters per line. Keep UI labels, names, and numbers together.
4. Time to the mouth. A cue appears as the line starts and clears as it ends. Minimum ~1 second on screen, maximum ~7 seconds. Split anything longer. Reading speed ≤ ~17 characters/second including spaces. If a cue fails the speed test, split it or extend the out-time into a real pause — never into the next speaker's words.
5. Clean mode: drop um / uh / false starts unless they carry meaning. Keep names, numbers, and UI terms exact. Verbatim mode: keep the fillers; still obey line length and speed.
6. Fill `templates/cue-sheet.md`. One row per cue: in, out, text, char counts, chars/sec, pass/fail. Fix fails before export. Cues must not overlap.
7. Fill `templates/export.md` with a valid SRT or WebVTT body plus a short QA list (sync, speed, overlap, orphans, speaker labels). Hand burn-in and motion type to motion-craft. Hand the spoken script itself to beat-script or hook-3s.
8. If you have no times and no duration, ask at most two questions, then either wait or export a cue list with every clock marked `[TIME: guess]`. Never fake a confident clock.

## Output format

- Source card (mode, format, duration, time evidence, labeled guesses)
- Cue sheet (timed rows with length and speed checks)
- Export card (valid SRT or WebVTT plus QA)
- One line: caption track; not burn-in, not a spoken script, not a social caption pack

## Safety

- Follow `references/cue-rails.md` with no exceptions.
- No phishing, malware pitches, fake medical cures, or guaranteed-return finance scripts.
- No CSAM or sexual content involving minors.
- Do not invent words the speaker did not say (clean-up of fillers is the one allowed edit in clean mode).
- Do not invent speaker names, timestamps presented as fact, or a duration you were not given.

## Stop conditions

- Three templates delivered, or refused for burn-in/kinetic type (motion-craft) / spoken-script writing (beat-script or hook-3s) / thumbnail (thumb-pack) / assembly (shot-to-cut) / social-post copy / scam / missing speech with no way to proceed
Version History
Comments (0)
No comments yet. Be the first!
Sign in to leave a comment.