# Notification: shorter video prompts

Prompt revision only. No images, video clips or audio have been regenerated. The current Remotion preview and exports still use the previous timing.

## Timing comes from the audio

The 18 existing readings total 106.208 seconds. The revised scene plan is 113.133 seconds, plus a two-second power-off ending, for 115.133 seconds overall. The old scene plan was 180 seconds.

Reels start their reading at 0.10 seconds and leave 0.10 seconds after it, rounded up to the next 30 fps frame. Lock-screen readings start at 0.40 seconds to allow the notification to arrive, with the same 0.10-second tail. Scroll transitions overlap the opening picture; they add no hold after the reading. Audio stays at its original speed and retains every word.

The seven future video requests total 62 whole seconds instead of 79. Their edited lengths follow the audio exactly. Trim only surplus picture from whole-second generation rounding. Do not stretch short clips, slow down the voice, or pad the timeline to the generated file length. Frame 13 retains roughly thirteen seconds because its reading alone is 12.725 seconds.

| Frame | Existing audio | Proposed edit | Future video request |
|---|---:|---:|---:|
| 01 | 3.855 s | 4.367 s | Remotion only |
| 02 | 3.994 s | 4.500 s | Remotion only |
| 03 | 9.474 s | 9.700 s | 10 s |
| 04 | 6.084 s | 6.300 s | 7 s |
| 05 | 6.177 s | 6.400 s | 7 s |
| 06 | 3.669 s | 3.900 s | Remotion only |
| 07 | 8.499 s | 9.000 s | Remotion only |
| 08 | 4.969 s | 5.500 s | Remotion only |
| 09 | 3.019 s | 3.533 s | Remotion only |
| 10 | 8.173 s | 8.400 s | 9 s |
| 11 | 7.245 s | 7.467 s | 8 s |
| 12 | 7.152 s | 7.367 s | 8 s |
| 13 | 12.725 s | 12.933 s | 13 s |
| 14 | 2.508 s | 3.033 s | Remotion only |
| 15 | 6.594 s | 7.100 s | Remotion only |
| 16 | 2.601 s | 3.133 s | Remotion only |
| 17 | 6.130 s | 6.633 s | Remotion only |
| 18 | 3.344 s | 3.867 s | Remotion only |

## Order of work

1. Keep the existing reference images and approved character identities.
2. Measure and verify all existing poem recordings first. This is complete; measured lengths and audio hashes are saved in plan.json.
3. Review the prompts below. They are drafts only, with no new requests submitted.
4. When generation resumes, use revised-clip-requests.json and those measured durations. Keep the original batch records unchanged.
5. Trim each picture to its proposed edit duration. Use the original recording at normal speed; never extend a scene to fill a video file.
6. Preserve the scrolls through 03, 04, 05, 06 and 10, 11, 12, 13; lock transitions into 07 and 14; moving cards on the same wallpaper for 07, 08, 09 and 14 through 18. After the final line, fade the entire phone to black.

## Frame 06 sound annotation

If an instrumental is added later, use only the 3.900-second frame-06 slot under its exact existing reading. Quick fade in and out; no competing voice. Reference only, not embedded.

Existing mood reference: [snowfall by Øneheart & reidenshi](https://www.youtube.com/watch?v=U1m46getoEw). No track has been added.

## Revised Seedance prompts

### Frame 03: Rich son, poor son

Audio: 9.474 s. Edit: 9.700 s. Future generation: 10 s.

Existing reference: `production/references/frame-03-clean.png`

Banana family A. Match these same father and sons when they return in frame 10.

Generate 10 seconds of picture. Final edit uses only the first 9.700 seconds for an existing 9.474-second reading, starting at 0.10s. Do not stretch the reading or add a silent hold to use the whole generated file. Any rounded duration surplus is discarded.

Open already mid-argument with the father pointing at the navy-suited rich son, who instantly smirks. At 2.39s snap his pointing hand toward the hoodie-wearing poor son. At 4.73s the poor son throws one incredulous palm upward. At 6.95s the rich son bursts into a mocking laugh while his brother glares. One readable reaction per spoken beat; no waiting between gestures. Cut away immediately after the last spoken word.

Vertical 9:16 silent image-to-video. Use the existing supplied image unchanged as the identity and composition reference. Immediate action from the first frame, brisk natural gestures, clear reactions timed to the listed narration cues, subtle handheld smartphone movement. One continuous shot with a restrained quick push-in on its strongest reaction; no slow motion, establishing shot, fades or outro. Keep faces and hands inside the central crop for a 393:852 iPhone screen. Maintain faces, wardrobe, props and background. No phone frame, status bar, app UI, notifications, captions, generated text, logo or watermark. No extra characters, morphing or wardrobe changes. Generate no speech, music or incidental dialogue. The existing ElevenLabs recording supplies every spoken word verbatim; do not invent or paraphrase any words.

Exact existing voiceover:

```text
You son: are going to grow up rich son
and you son: will grow up poor.
"Daddy why can't I have anything?"
hahahah, look at you brother, so poor!
```

### Frame 04: Twenty thousand

Audio: 6.084 s. Edit: 6.300 s. Future generation: 7 s.

Existing reference: `frames/frame-04-influencer-v3.png`

Human influencer B, Asian man, 24, sleeve tattoos on both arms, sunglasses, cigar. This is a human replacement for the orange character.

Generate 7 seconds of picture. Final edit uses only the first 6.300 seconds for an existing 6.084-second reading, starting at 0.10s. Do not stretch the reading or add a silent hold to use the whole generated file. Any rounded duration surplus is discarded.

Open with the 24-year-old Asian human influencer already leaning toward the selfie lens and pointing with his free hand. At 2.02s he opens that palm to emphasize starting from zero. At 3.71s he jabs toward the lens once, then gives a knowing smirk at 4.73s. Hold a smoking cigar in the other hand with a small curl of smoke throughout; omit the slow draw-and-exhale sequence. Preserve sunglasses, black short-sleeved shirt, gold watch and full sleeve tattoos on both arms. Use a brisk natural sales pitch gesture, then cut immediately after the final line.

Vertical 9:16 silent image-to-video. Use the existing supplied image unchanged as the identity and composition reference. Immediate action from the first frame, brisk natural gestures, clear reactions timed to the listed narration cues, subtle handheld smartphone movement. One continuous shot with a restrained quick push-in on its strongest reaction; no slow motion, establishing shot, fades or outro. Keep faces and hands inside the central crop for a 393:852 iPhone screen. Maintain faces, wardrobe, props and background. No phone frame, status bar, app UI, notifications, captions, generated text, logo or watermark. No extra characters, morphing or wardrobe changes. Generate no speech, music or incidental dialogue. The existing ElevenLabs recording supplies every spoken word verbatim; do not invent or paraphrase any words.

Exact existing voiceover:

```text
How I'd make 20k per month
starting from zero
here's how I'd do it
in 50 seconds
```

### Frame 05: Selfie confession, part one

Audio: 6.177 s. Edit: 6.400 s. Future generation: 7 s.

Existing reference: `frames/frame-05-11-brunch-selfie-v4.png`

Selfie character C: the mixed-race woman previously shown in frame 12, warm medium-brown skin, brown eyes, curly dark brunette ponytail, floral brunch dress. Frame 11 reuses this exact portrait.

Generate 7 seconds of picture. Final edit uses only the first 6.400 seconds for an existing 6.177-second reading, starting at 0.10s. Do not stretch the reading or add a silent hold to use the whole generated file. Any rounded duration surplus is discarded.

Start mid-blush stroke, face filling the front-facing selfie camera. At 1.25s her eyes flick straight into the lens; at 2.42s stop the brush and raise an eyebrow at the engagement confession. At 4.81s point once with the brush and lean slightly closer. Cut on the unfinished confession after the recorded line ends. Same adult mixed-race woman with warm medium-brown skin, brown eyes, curly dark brunette ponytail and light floral brunch dress. Show the dress neckline, not a mirror or reflected phone. No introductory posing or slow setup.

Vertical 9:16 silent image-to-video. Use the existing supplied image unchanged as the identity and composition reference. Immediate action from the first frame, brisk natural gestures, clear reactions timed to the listed narration cues, subtle handheld smartphone movement. One continuous shot with a restrained quick push-in on its strongest reaction; no slow motion, establishing shot, fades or outro. Keep faces and hands inside the central crop for a 393:852 iPhone screen. Maintain faces, wardrobe, props and background. No phone frame, status bar, app UI, notifications, captions, generated text, logo or watermark. No extra characters, morphing or wardrobe changes. Generate no speech, music or incidental dialogue. The existing ElevenLabs recording supplies every spoken word verbatim; do not invent or paraphrase any words.

Exact existing voiceover:

```text
Get ready with me
while I talk about how
I almost ruined my engagement
so my fiance and I
```

### Frame 10: Rich brother returns

Audio: 8.173 s. Edit: 8.400 s. Future generation: 9 s.

Existing reference: `production/references/frame-10-clean.png`

Banana family A from frame 03; keep the faces and room consistent. The poorer son now wears the ivory suit.

Generate 9 seconds of picture. Final edit uses only the first 8.400 seconds for an existing 8.173-second reading, starting at 0.10s. Do not stretch the reading or add a silent hold to use the whole generated file. Any rounded duration surplus is discarded.

Open on the navy-suited banana brother already wide-eyed at the ivory-suited brother. At 2.23s he points at him in disbelief. At 3.66s the ivory-suited brother snaps a single stable stack of money into view, prompting a short recoil. At 5.93s the ivory-suited brother leans forward with a proud grin and one confident gesture. Cut immediately after the hard-work line. Preserve the banana family faces, stems and room from frame 03. No extra money reveal or empty reaction hold.

Vertical 9:16 silent image-to-video. Use the existing supplied image unchanged as the identity and composition reference. Immediate action from the first frame, brisk natural gestures, clear reactions timed to the listed narration cues, subtle handheld smartphone movement. One continuous shot with a restrained quick push-in on its strongest reaction; no slow motion, establishing shot, fades or outro. Keep faces and hands inside the central crop for a 393:852 iPhone screen. Maintain faces, wardrobe, props and background. No phone frame, status bar, app UI, notifications, captions, generated text, logo or watermark. No extra characters, morphing or wardrobe changes. Generate no speech, music or incidental dialogue. The existing ElevenLabs recording supplies every spoken word verbatim; do not invent or paraphrase any words.

Exact existing voiceover:

```text
Wait, it's you?
Brother, you were poor.
No way you have all this money.
I learned the value of hard work
```

### Frame 11: Selfie confession, part two

Audio: 7.245 s. Edit: 7.467 s. Future generation: 8 s.

Existing reference: `frames/frame-05-11-brunch-selfie-v4.png`

Exactly character C from frame 05: same mixed-race adult woman, face, brown eyes, curly brunette ponytail, floral brunch dress, brush, room, lighting and selfie crop.

Generate 8 seconds of picture. Final edit uses only the first 7.467 seconds for an existing 7.245-second reading, starting at 0.10s. Do not stretch the reading or add a silent hold to use the whole generated file. Any rounded duration surplus is discarded.

Same exact woman, face, brunette curly ponytail, floral brunch dress, brush, room, lighting and front-camera crop as frame 05. Start with the brush raised and a knowing look directly into the lens. At 1.06s give one tiny nod, then at 3.12s stop applying makeup and lift an eyebrow. At 4.66s point once with the brush and lean in with an incredulous expression. Cut as the unfinished recorded confession ends. No mirror, outfit change or repeated establishing shot.

Vertical 9:16 silent image-to-video. Use the existing supplied image unchanged as the identity and composition reference. Immediate action from the first frame, brisk natural gestures, clear reactions timed to the listed narration cues, subtle handheld smartphone movement. One continuous shot with a restrained quick push-in on its strongest reaction; no slow motion, establishing shot, fades or outro. Keep faces and hands inside the central crop for a 393:852 iPhone screen. Maintain faces, wardrobe, props and background. No phone frame, status bar, app UI, notifications, captions, generated text, logo or watermark. No extra characters, morphing or wardrobe changes. Generate no speech, music or incidental dialogue. The existing ElevenLabs recording supplies every spoken word verbatim; do not invent or paraphrase any words.

Exact existing voiceover:

```text
Part two,
since y'all kept asking.
So after I caught him cheating
on me with my best friend I
```

### Frame 12: Day fifty: van-life update

Audio: 7.152 s. Edit: 7.367 s. Future generation: 8 s.

Existing reference: `frames/frame-12-brunette-traveler-v4.png`

Traveler D: petite white adult woman, age 25, fair skin, blue eyes and dark brunette ponytail, with the delicate facial features of the former frame-11 character. Keep the forest-green leggings and matching crop top.

Generate 8 seconds of picture. Final edit uses only the first 7.367 seconds for an existing 7.152-second reading, starting at 0.10s. Do not stretch the reading or add a silent hold to use the whole generated file. Any rounded duration surplus is discarded.

The camper-van door is already open on the first frame. The petite white adult brunette woman with a ponytail immediately turns toward the camera, preserving her forest-green leggings and matching crop top. At 1.82s give a quick shrug; at 2.98s gesture once toward the unfinished van interior. At 4.45s point out toward the valley and grin at the camera. A small breeze moves her ponytail. Cut right after day fifty. No slow door-opening reveal, landscape pan or lingering pose.

Vertical 9:16 silent image-to-video. Use the existing supplied image unchanged as the identity and composition reference. Immediate action from the first frame, brisk natural gestures, clear reactions timed to the listed narration cues, subtle handheld smartphone movement. One continuous shot with a restrained quick push-in on its strongest reaction; no slow motion, establishing shot, fades or outro. Keep faces and hands inside the central crop for a 393:852 iPhone screen. Maintain faces, wardrobe, props and background. No phone frame, status bar, app UI, notifications, captions, generated text, logo or watermark. No extra characters, morphing or wardrobe changes. Generate no speech, music or incidental dialogue. The existing ElevenLabs recording supplies every spoken word verbatim; do not invent or paraphrase any words.

Exact existing voiceover:

```text
I just quit my 9-5
with no backup plan
to renovate a van and
travel the world: day 50.
```

### Frame 13: Still running

Audio: 12.725 s. Edit: 12.933 s. Future generation: 13 s.

Existing reference: `production/references/frame-13-clean.png`

Fruit fitness pair E. Preserve the strawberry bodybuilder and banana runner from the reference.

Generate 13 seconds of picture. Final edit uses only the first 12.933 seconds for an existing 12.725-second reading, starting at 0.10s. Do not stretch the reading or add a silent hold to use the whole generated file. Any rounded duration surplus is discarded.

The strawberry bodybuilder is already flexing and the banana is already running on the treadmill on the first frame. At 2.49s the strawberry points at the running banana; at 5.00s both briefly glance toward the lens. At 6.61s the strawberry breaks into a smug laugh and hits one bigger flex. The banana answers with one exasperated side-eye while maintaining a coherent stride. Keep the final spoken taunt visually active with small natural reactions; no repeated flex routine or dead hold. End immediately after the running question. Maintain the original fruit identities, treadmill contacts and gym. This passage needs about thirteen seconds to retain every poem word; make its actions brisk without shortening the narration.

Vertical 9:16 silent image-to-video. Use the existing supplied image unchanged as the identity and composition reference. Immediate action from the first frame, brisk natural gestures, clear reactions timed to the listed narration cues, subtle handheld smartphone movement. One continuous shot with a restrained quick push-in on its strongest reaction; no slow motion, establishing shot, fades or outro. Keep faces and hands inside the central crop for a 393:852 iPhone screen. Maintain faces, wardrobe, props and background. No phone frame, status bar, app UI, notifications, captions, generated text, logo or watermark. No extra characters, morphing or wardrobe changes. Generate no speech, music or incidental dialogue. The existing ElevenLabs recording supplies every spoken word verbatim; do not invent or paraphrase any words.

Exact existing voiceover:

```text
You will get in shape with steroids and
you will get in shape by running
Let's see who looks better
hahahah, look at me! I'm fit, and you're still running?
```
