Back to Blog
Production

Directing Dialogue and Acting: How to Prompt for Character Speech, Pauses & Emotional Beats

Learn how LTX-2.3's new gated-attention text connector changes dialogue prompting with the segmentation pattern, acting cues, and worked examples.

LTX Team
Production
Key Takeaways
  • LTX-2.3's redesigned text connector reads acting direction at the phrase level, so segmenting dialogue into short phrases + cues (instead of full quoted sentences) is now the key prompting pattern.
  • Use four cue types — eye-line/gaze, pauses/pacing, voice quality, physical beats — one per beat, plus explicit audio direction (acoustic space, voice texture, ambient sound) at the end of the prompt.
  • Old LTX-2 prompts need re-segmenting for LTX-2.3: long unbroken sentences and vague emotional adjectives now perform worse since the model no longer ignores unstructured direction.

LTX-2.3 follows acting direction at the beat level. Break a line of dialogue into short phrases. Slot acting cues between them: a pause, an eye-line shift, a voice that cracks. The model renders the performance, not just the words.

The unlock is a redesigned text connector with gated attention that reads granular direction more literally than LTX-2 did. The cost is that long unbroken sentences and vague emotional adjectives now perform worse than they used to, because the model is listening more carefully to the structure you give it.

This guide is for creators writing prompts for speaking characters on LTX-2.3: short films, dialogue-driven ads, scripted shorts, voice-led product demos.

It covers why the model handles dialogue differently than 2.x, the segmentation pattern that unlocks the new behavior, the acting-direction vocabulary the model reliably responds to, how to pair dialogue with audio descriptions, and three worked examples you can adapt directly.

Why LTX-2.3 handles dialogue better

LTX-2 already used Gemma 3 as its text-encoding backbone. LTX-2.3 keeps that backbone but quadruples the size of the text encoder and adds a new gated-attention text connector between the encoder and the video transformer.

The practical consequence is tighter prompt adherence: the model maps specific words in the prompt to specific frames in the output more reliably than the LTX-2 baseline.

For dialogue prompts, that means timing, pacing, and emotional cues that LTX-2 would smear across the clip now land on the intended beat.

Two implications follow. First, prompt length matters more: longer, more descriptive prompts consistently outperform short ones on LTX-2.3, because the model has more anchors to align the generation against.

Second, LTX-2 prompts often need to be rewritten for LTX-2.3; the model is now sensitive to direction that LTX-2 would silently ignore, and that includes both deliberate cues and accidental noise.

The segmentation pattern

The pattern: write one phrase of dialogue at a time, then a short acting direction, then the next phrase. Don't quote a full sentence. Don't separate the dialogue from the acting into different paragraphs. Interleave them in the order the camera would see them.

What does a segmented prompt look like?

Compare an unsegmented prompt against a segmented one for the same line.

Unsegmented (LTX-2 style):

A middle-aged man with greying hair speaks slowly: "I remember after you kids came along, your mom said something to me I never quite understood." He looks sad. The camera zooms into his face.

Segmented (LTX-2.3 style):

A middle-aged man with greying hair speaks in a sad, slow-paced voice, "I remember after you kids came along..." He pauses and looks to the side, then continues, "your mom..." His eyes widen momentarily. He finishes with a cracking voice, "said something to me I never quite understood." The camera slowly zooms into his face. The audio is crisp with faint room tone.

The segmented version gives the model explicit beats: the opening phrase, a directed pause, a one-word fragment, a micro-expression, a voice change, the closing phrase, and a camera move. LTX-2.3 renders each beat at the right moment in the clip; the unsegmented version leaves the model to guess where to put the pauses.

Why long unbroken sentences fail

LTX-2.3 will still try to perform a long unbroken line, but the timing collapses. The model rushes the delivery to fit the duration, or fills the gap with idle motion that contradicts the line. Segmentation gives the model permission to slow down, hold the frame, and let the beat land.

How do you write acting direction?

The model reliably responds to four families of direction. Use them sparingly, one cue per beat is enough, and keep the language specific.

Eye-line and gaze

  • "looks to the side"
  • "glances down"
  • "eyes widen momentarily"
  • "holds her gaze on him"
  • "looks past the camera"

Gaze cues anchor the character's attention and break up otherwise static talking-head shots. They work best between phrases, not during them.

Pauses and pacing

  • "pauses, then continues"
  • "trails off"
  • "holds the silence for a beat"
  • "finishes quietly"
  • "speaks slowly, weighing each word"

Pacing cues are the single highest-leverage direction in a dialogue prompt. A pause inside a line creates a moment the audience reads as deliberation; without it, the line reads as a recitation.

Voice quality

  • "cracking voice"
  • "whispered, low energy"
  • "voice tight with restraint"
  • "warm, conversational delivery"
  • "forced confidence"

Voice quality cues are the bridge between the dialogue and the audio prompt. LTX-2.3's audio generation is cleaner and less artifact-prone than LTX-2's, so voice descriptors translate to perceptible texture differences in the rendered audio.

Physical beats

  • "his shoulders drop"
  • "leans forward slightly"
  • "his hand tightens on the edge of the table"
  • "a half-smile fades"
  • "swallows before continuing"

Physical beats reinforce emotional shifts without requiring the actor's face to do all the work. They also give the model something concrete to ground the frame in during pauses.

How should you pair dialogue with audio?

LTX-2.3 generates audio alongside video with fewer artifacts, cleaner room tone, and less over-processed compression than LTX-2. That makes audio descriptions in the prompt more impactful: when you describe the acoustic environment, the model can actually deliver it.

Audio direction worth including

  • Acoustic space — "faint room tone," "echo of a marble lobby," "outdoor wind with distant traffic"
  • Voice texture — "low-mic warmth," "slightly breathy," "the resonance of a small studio"
  • Silence shape — "quiet between phrases," "long silence after the line," "the room holds the silence"
  • Ambient cues — "a clock ticks softly in the background," "rain on glass behind him"

Audio direction goes at the end of the prompt, after the dialogue and acting beats are laid out. The model treats it as the environmental wrapper around the performance.

How do you direct the camera around dialogue?

Camera movement in a dialogue shot is part of the acting direction. The cues that work best on LTX-2.3 are slow, motivated movements that underline a beat rather than competing with it.

  • Slow push-ins on the emotional turn of a line
  • Hold the frame during pauses: the camera goes still, the actor carries the beat
  • Slight rack focus from background to face when the character shifts attention
  • Avoid wide horizontal pans during dialogue: the model uses them to mask static moments, which reads as bored coverage

Camera direction sits at the same prompt level as physical beats, interleaved between dialogue phrases, not stacked into a separate paragraph.

Worked examples

Three prompts you can adapt. Each is segmented, beat-anchored, and ends with audio direction.

A regretful father monologue

A man in his late fifties with greying stubble sits at a wooden kitchen table in a warm, dim room. He speaks in a slow, weary voice, "I never told you this..." He pauses and looks at his hands, then continues, "but I almost left..." His shoulders drop. He finishes, voice catching, "the year you were born." The camera holds the frame. Faint room tone, a refrigerator hum, no music.

A confident product founder pitch

A woman in her early thirties stands in front of a clean white backdrop, wearing a black turtleneck. She speaks with warm, conversational confidence, "We didn't set out to build another tool..." She pauses, makes eye contact with the camera, then continues with quiet certainty, "we built the one we wanted to use." A half-smile lands at the end of the line. The camera is locked off in a medium-close-up. The audio is crisp, low-mic warmth, soft studio ambience.

A frightened child whispering

A child of about eight sits on the floor in a darkened bedroom, back against the wall, a flashlight held in her lap. She speaks in a tight whisper, "I can hear it again..." She holds her breath for a beat, then continues, voice trembling, "in the closet." Her eyes widen. The camera slowly pushes in. Audio: very quiet room tone, the faint creak of a floorboard somewhere off-frame, no music.

What are the common mistakes?

Four patterns account for most weak dialogue results on LTX-2.3.

Writing dialogue as one block of quoted text. LTX-2.3 will perform it, but it will perform it as a recitation. The fix is segmentation: break the line at every natural beat and insert a direction.

Over-directing every micro-beat. If every phrase carries three cues, the model has nothing to prioritize and the performance reads as twitchy. One direction per beat is the working limit.

Ignoring the audio prompt. LTX-2.3's audio quality is good enough to carry emotion on its own, but only if you describe what you want. A prompt without any audio direction leaves the model to default to a flat studio sound.

Using LTX-2 prompts unchanged on LTX-2.3. Prompts that worked on LTX-2 because the model ignored half of them will perform differently on LTX-2.3, which doesn't ignore the noise. Review and trim before re-using.

Summary

LTX-2.3's gated-attention text connector follows acting direction at the phrase level, which makes the segmentation pattern — one phrase of dialogue, one acting cue, the next phrase — the working pattern for prompting speaking characters. Use the four direction families (eye-line, pacing, voice quality, physical beats), one cue per beat, pair the dialogue with explicit audio direction, and let the camera underline rather than compete with the performance. Prompts that worked on LTX-2 will need to be re-segmented before they perform well on LTX-2.3 because the model now listens to structure that 2.x silently ignored.

Start prompting on LTX-2.3 today via the LTX API or download the open weights on HuggingFace.

Table of Contents: