- Foley LoRA for LTX-2.3-22B generates synchronized Foley and sound effects directly from silent video, covering impacts, footsteps, materials, machines, and weather, with no speech or music.
- Works video-to-audio (adds sound to a muted clip) and text-to-audio (generates sound from a description alone), keeping the input video fixed while only audio latents are generated.
- Recommended setup: LoRA strength 0.8–1.0, guidance 6.0, 30 steps, 960×544 @ 24fps, and the step-500 checkpoint for best loudness.
AI-generated video looks stunning, but it usually arrives without sound. The Foley LoRA for LTX-2.3-22B changes that by generating synchronized, realistic sound effects directly from silent video.
Whether you need footsteps, impacts, weather, water, or machines, this LoRA adds the audio layer that makes a scene feel real. No speech, no music bed — just Foley.
See It In Action
These examples show the Foley LoRA applied to a range of actions and materials. Each clip is generated from a text prompt that describes the visible source, action, and material.
Fireball explosion: Deep boom, concussive blast, roaring flames.
Glass smash: Heavy whoosh, sharp shatter, tinkling shards.
Hydraulic press: Heavy hydraulic whirr, metal crunch, mechanical clank.
Snow footsteps: Icy snow cracking, crisp crunchy footfalls.
Rain close-up: Delicate drops, soft splashes, steady rain hiss.
Factory robot: Servo whirrs, metallic clicks, pneumatic hisses.
Side-by-side comparisons show the same real footage muted (left) and with Foley LoRA audio (right).
Blacksmith hammering — muted vs. Foley LoRA.
Wet slush footsteps — muted vs. Foley LoRA.
Wooden boardwalk footsteps — muted vs. Foley LoRA.
What It Does
The Foley LoRA is a lightweight adapter for LTX-2.3-22B. It takes a silent video and generates a synchronized sound track of the actions you would actually hear.
It is video-to-audio first. Feed it a muted clip and it returns the same clip with matching sound. It also works for text-to-audio Foley if you want to generate sound from a description alone.
It does not add speech or music. It focuses on Foley and SFX — the layer of sound that grounds a scene in reality.
Covered material families:
- Impacts & breaks: glass shatter, ceramic break, wood snap, metal hammer on anvil
- Footsteps & surfaces: snow crunch, wet slush, gravel, dry leaves, wooden boardwalk
- Materials & handling: rubber squeak, padded boxing gloves, keyboard clicks, cloth rustle
- Tools & machines: hydraulic press, metal-cutting machine, robot arm, chainsaw, arc welder
- Weather & water: rain on a roof, puddle splashes, wave crash, dripping water
- Ambient events: campfire, espresso pour, cathedral bell, sword draw
How It Works
The LoRA targets the audio attention blocks and the video-to-audio cross-attention in LTX-2.3-22B. During inference, the input video is kept fixed, and only the audio latents are generated.
Training used roughly 5,374 short clips at 960×544, 24 fps, about 3.7 seconds each. The dataset combined a cleaned Foley rebuild with the FoleyBench academic dataset. Captions describe the visible action and explicitly exclude speech and music.
The model trained for 3,000 steps, but the step-500 checkpoint is recommended for publication work because later checkpoints tend to over-suppress volume.
When to Use It
Use the Foley LoRA whenever you have a silent or muted clip that needs a sound bed. It is ideal for:
- Adding sound to AI-generated video from LTX-2.3 or other pipelines
- Replacing unwanted original audio with clean Foley
- Prototyping sound design quickly without sourcing stock libraries
- Generating material-specific sounds where the visual contact is clear
It is not a substitute for dialogue synthesis, music generation, or final professional mixing. Use it for Foley and SFX layers.
Get Started
Recommended inference recipe:
- Base model: LTX-2.3-22B
- LoRA strength: 0.8–1.0
- Guidance scale: 6.0
- Inference steps: 30
- Resolution: 960×544 @ 24 fps
- Frame count: 1 + 8×n (89 frames is the main training bucket; tested up to 169)
- STG: scale 1.0, block [29], mode stg_av
- Checkpoint: step 500 (best loudness)
Prompting tip:
Write one to three dense sentences that name the action, material, acoustic character, and timing. End with: No speech is present. No music is present.
Example:
A sledgehammer smashes a glass bottle: a heavy whoosh, a sharp glass shatter and tinkling shards. No speech is present. No music is present.
Base negative prompt:
music, melody, song, singing, vocals, score, soundtrack, speech, dialogue, talking, narration, muffled, dull, distant, low-pass filtered, underwater, distorted, clipped, static, room tone only, silence, generic sound effect
Get the model and workflow:
Model: Lightricks/LTX-2.3-22b-LoRA-Foley-V2A on Hugging Face
ComfyUI nodes: ComfyUI-LTXVideo on GitHub
Base model: Lightricks/LTX-2.3 on Hugging Face
Credits
Demo footage sources:
- Side-by-side comparison clips are from Pexels under the Pexels Free License.
- Hero example clips were generated end-to-end with no stock footage: Seedream v5 Pro (text → still) → LTX-2.3 I2V (still → muted video) → Foley LoRA v4 (video → audio).
- Publishing assets, ComfyUI workflows, and the model card are included in the Hugging Face release.
Acknowledgments:
- Base model: Lightricks/LTX-2.3
- Training infrastructure: LTX-2 Community Trainer
- Dataset: FoleyBench (CC BY-NC-SA) and a curated Foley rebuild set
License: LTX-2 Community License. Dataset components may carry additional terms.
