- Lead crowd prompts with shot type and environment, not character counts — the model can't "count" extras, but a precisely named location (Tokyo izakaya, Mumbai market) implies the right crowd density and behavior.
- Use qualitative density words (sparse, packed, shoulder-to-shoulder) anchored to the setting, one camera move per generation, and active camera motion paired with high-energy crowds to avoid smeared faces.
- For broadcast-grade control, layer the Pose-Control IC-LoRA (built for LTX-2 19b) on top of prompting — but confirm it matches your base checkpoint, since the Union-Control adapter targets LTX-2.3 22b instead.
Crowd shots are the hardest test for any AI video model. A protagonist close-up has one face to render correctly; a crowded train station has fifty, all moving, all needing to look like distinct people in coordinated wardrobes under one consistent light source.
LTX-2.3 can produce credible crowd shots, but only when the prompt leads with shot framing and environment, not with character counts.
This guide covers the prompt patterns that work for LTX-2.3 crowd generation, the scenario templates you can lift directly into your own briefs, and the IC-LoRA controls that move crowds from "AI-looking" to broadcast-ready.
Prerequisites: A working LTX setup via the self-hosted open-source pipeline, or hosted API access to a current LTX model.
Why are crowd shots so hard?
Crowd shots stress the model on three independent axes at once: identity diversity across many faces, spatial coherence as bodies overlap and occlude each other, and temporal consistency as those bodies move through the scene. A single failure on any axis breaks the illusion. You see it as duplicate faces in the third row, as a hand clipping through a shoulder, or as a passerby who teleports between frames.
LTX-2.3's documented prompting guidance gives the workable angle on this. The repository's prompting section recommends "detailed, chronological descriptions of actions and scenes" within a single flowing paragraph, with literal and precise descriptions and a 200-word ceiling.
That framing favors describing the environment and the camera, then letting the model populate the crowd, rather than enumerating extras one by one.
LTX-2.3 crowd-prompting fundamentals
LTX-2.3 responds best to prompts that establish what the camera is doing first, then describe the environment in cinematographer's terms, and only then mention the crowd as a property of that environment.
The documented prompt structure (main action, then movements, then appearances, then environment, then camera angles, then lighting) translates almost directly into a crowd-prompt template.
Lead with shot type, not character count
A prompt that opens with "A wide shot of a hundred extras" front-loads a number the model cannot reliably count to. A prompt that opens with "A high-angle wide shot of a packed European train platform during morning rush hour" gives the model a shot type, a location, and a density cue without ever asking it to count. The model populates the platform with the right kind of bodies, in roughly the right density, doing roughly the right things.
Describe the environment, let the model populate it
Specific environments imply specific crowds. A "Tokyo izakaya at 11pm" produces salarymen, neon signs, and small clusters of two and three. A "Mumbai street market at golden hour" produces dense foot traffic, fruit stalls, and movement on multiple depth planes. The environment is doing most of the work; the prompt's job is to name the place precisely enough that the model can infer the crowd.
Density modifiers that actually work
Avoid numeric crowd counts ("about 50 people"). Use qualitative density words that map cleanly onto visual references: sparse, scattered, steady foot traffic, busy, packed, shoulder-to-shoulder. Anchor those to the environment ("packed concert audience" reads differently from "packed subway car") so the model has a visual prior.
Prompt patterns by scenario
Stadium and arena crowds
Stadium crowds are the easiest crowd shot because they have a strong geometric structure (rows, sections, uniform direction of attention). Lead with the camera position, then the venue, then the energy, then lighting.
Example prompt: "A high-angle wide shot from the upper deck of a packed soccer stadium during a night match, fans in red and white scarves moving in coordinated waves, stadium floodlights overhead, the pitch visible in the lower third of the frame, slow camera pan from left to right, atmospheric haze from flares in the lower stands."
Concert and festival audiences
Concert audiences need motion. A static crowd at a concert reads as wrong instantly. Specify the energy and reference a known music genre to set the right body language.
Example prompt: "A medium-wide shot from front of stage at an outdoor electronic music festival at dusk, the audience packed shoulder to shoulder with hands raised, bodies swaying to the beat, stage lights cutting through fog behind the camera, the silhouette of the crowd against a deep purple sky, handheld camera with slight motion."
Protest and rally scenes
Protest scenes carry political weight, so accuracy and neutrality matter. Describe the visual elements and avoid loaded specifics. Lead with the geography of the gathering.
Example prompt: "A wide shot down a city avenue filled with marchers carrying handmade signs, mid-morning overcast light, signs and banners varied in color, marchers in winter coats, the crowd extending toward a vanishing point in the distance, slow tracking camera moving forward through the crowd, ambient noise of chanting."
Busy street and market scenes
Markets are great for testing depth and parallax. Multiple planes of activity, varied wardrobe, products on stalls, all running simultaneously. The prompt should specify foreground, middle ground, and background separately.
Example prompt: "A handheld medium shot moving through a Mumbai street market at golden hour, foreground vendors arranging mangoes on a wooden cart, middle ground steady foot traffic in colorful saris and shirts, background stalls hung with fabric in pinks and oranges, warm late-afternoon light cutting between buildings, ambient market sound."
Corporate event and conference rooms
Conference rooms break a lot of AI video models because the wardrobe is uniform (business attire), the lighting is even and unflattering, and the bodies are mostly stationary. Use camera movement to break the static feel.
Example prompt: "A slow dolly shot through a packed corporate conference auditorium during a keynote, attendees in business attire seated facing the stage, soft event lighting from above, stage backlit in blue, occasional movements as audience members shift in seats, faces lit by laptop screens in the back rows."
How do you control crowd appearance?
Wardrobe and time-period cues
Wardrobe is the fastest lever for crowd believability. A 1970s street market and a 2026 street market produce identical crowd geometries but read completely differently in trailers and ads. Specify era ("late 1980s commuters in trench coats and tailored suits") or aesthetic ("street style in muted earth tones") rather than listing garments.
Age, demographic, and energy descriptors
Avoid percentages or demographic counts. Use age-range words ("families with young children," "a crowd skewing twenty-something," "an older audience") and energy words ("attentive," "rowdy," "subdued"). LTX-2.3 follows these prompts more reliably than it follows explicit numerical breakdowns.
Lighting and mood for crowds
Crowd lighting is half of crowd believability. Strong directional light produces strong shadows and silhouettes, which read as cinematic. Flat overcast light kills depth and makes crowds look like collages. Name the light source ("backlit by a setting sun," "lit by sodium streetlamps," "natural overcast diffused light") to anchor the look.
Common crowd-prompt mistakes
Over-specifying individual extras
"A woman in a red dress in the third row, next to a man in a navy suit" forces the model to count and arrange. It rarely produces what you described. Describe the dominant wardrobe palette and the density, and let the model populate.
Camera-angle confusion
Mixing camera moves in the same prompt ("a wide shot that pans and zooms in") creates intra-shot inconsistency. Pick one camera move per generation. If you need a sequence, generate multiple clips and edit them.
Movement and motion-blur pitfalls
Asking for "fast motion" or "frenetic energy" with a static camera produces motion blur that smears faces and arms unpredictably. Pair high-energy crowd descriptions with active camera moves (handheld, dolly, tracking) so the motion is distributed across the frame instead of concentrated on each extra.
Before-and-after example prompts
Stadium prompt: weak to strong
Weak: "A crowded football stadium with 50,000 fans cheering."
Strong: "A high-angle wide shot of a packed European football stadium during a Champions League night match, fans in red and yellow waving scarves overhead, stadium floodlights from above creating long shadows, atmospheric haze drifting across the lower stands, slow camera pan across the home crowd, ambient roar of the stadium."
Street prompt: weak to strong
Weak: "A busy street with many people walking."
Strong: "A medium-wide shot down a rain-slicked Tokyo street at night, steady foot traffic of commuters in dark coats and clear umbrellas, neon signage reflected on the wet pavement, red and pink lights from convenience stores cutting through the rain, slow dolly forward at eye level, ambient sound of footsteps and distant traffic."
When should you use IC-LoRA alongside prompts?
For crowd shots specifically, providing a reference video or a pose track means you control the high-level geometry of the crowd while the model fills in wardrobe, lighting, and identity. This is the difference between hoping the model produces a march and dictating the march while the model dresses the marchers. Note that running multiple IC-LoRA groups simultaneously is documented as memory-intensive, so keep only one active per generation. Confirm your control adapter matches your base checkpoint version before building the workflow.
Summary
LTX-2.3 generates believable crowd shots when prompts lead with shot type and environment rather than character counts, use qualitative density words (sparse, packed, shoulder-to-shoulder) anchored to specific locations, and pair high-energy crowd descriptions with active camera moves to distribute motion across the frame.
For broadcast-grade crowd control, the Pose-Control IC-LoRA (built for the LTX-2 19b model) layered on top of prompt craft lets you dictate geometry while the model handles wardrobe and identity — the Union-Control adapter is built for LTX-2.3 (22b) and isn't a drop-in for an LTX-2 pipeline.
Start from the prompt templates above, run a single IC-LoRA group at a time, confirm it matches your base checkpoint version, and verify your durations against the documented frame-count constraints before deploying to production.