Back to Blog
Production

Bring Silent Videos to Life: Foley LoRA For LTX-2.3

Foley LoRA for LTX-2.3-22B generates synchronized, realistic sound effects directly from silent video — footsteps, impacts, weather, machines. No speech or music.

Ran Bensimon
Production
Custom Video Thumbnail Play Button
Key Takeaways
  • Foley LoRA for LTX-2.3-22B generates synchronized Foley and sound effects directly from silent video, covering impacts, footsteps, materials, machines, and weather, with no speech or music.
  • Works video-to-audio (adds sound to a muted clip) and text-to-audio (generates sound from a description alone), keeping the input video fixed while only audio latents are generated.
  • Recommended setup: LoRA strength 0.8–1.0, guidance 6.0, 30 steps, 960×544 @ 24fps, and the step-500 checkpoint for best loudness.

AI-generated video looks stunning, but it usually arrives without sound. The Foley LoRA for LTX-2.3-22B changes that by generating synchronized, realistic sound effects directly from silent video.

Whether you need footsteps, impacts, weather, water, or machines, this LoRA adds the audio layer that makes a scene feel real. No speech, no music bed — just Foley.

See It In Action

These examples show the Foley LoRA applied to a range of actions and materials. Each clip is generated from a text prompt that describes the visible source, action, and material.

Fireball explosion: Deep boom, concussive blast, roaring flames.

Glass smash: Heavy whoosh, sharp shatter, tinkling shards.

Hydraulic press: Heavy hydraulic whirr, metal crunch, mechanical clank.

Snow footsteps: Icy snow cracking, crisp crunchy footfalls.

Rain close-up: Delicate drops, soft splashes, steady rain hiss.

Factory robot: Servo whirrs, metallic clicks, pneumatic hisses.

Side-by-side comparisons show the same real footage muted (left) and with Foley LoRA audio (right).

Blacksmith hammering — muted vs. Foley LoRA.

Wet slush footsteps — muted vs. Foley LoRA.

Wooden boardwalk footsteps — muted vs. Foley LoRA.

What It Does

The Foley LoRA is a lightweight adapter for LTX-2.3-22B. It takes a silent video and generates a synchronized sound track of the actions you would actually hear.

It is video-to-audio first. Feed it a muted clip and it returns the same clip with matching sound. It also works for text-to-audio Foley if you want to generate sound from a description alone.

It does not add speech or music. It focuses on Foley and SFX — the layer of sound that grounds a scene in reality.

Covered material families:

  • Impacts & breaks: glass shatter, ceramic break, wood snap, metal hammer on anvil
  • Footsteps & surfaces: snow crunch, wet slush, gravel, dry leaves, wooden boardwalk
  • Materials & handling: rubber squeak, padded boxing gloves, keyboard clicks, cloth rustle
  • Tools & machines: hydraulic press, metal-cutting machine, robot arm, chainsaw, arc welder
  • Weather & water: rain on a roof, puddle splashes, wave crash, dripping water
  • Ambient events: campfire, espresso pour, cathedral bell, sword draw

How It Works

The LoRA targets the audio attention blocks and the video-to-audio cross-attention in LTX-2.3-22B. During inference, the input video is kept fixed, and only the audio latents are generated.

Training used roughly 5,374 short clips at 960×544, 24 fps, about 3.7 seconds each. The dataset combined a cleaned Foley rebuild with the FoleyBench academic dataset. Captions describe the visible action and explicitly exclude speech and music.

The model trained for 3,000 steps, but the step-500 checkpoint is recommended for publication work because later checkpoints tend to over-suppress volume.

When to Use It

Use the Foley LoRA whenever you have a silent or muted clip that needs a sound bed. It is ideal for:

  • Adding sound to AI-generated video from LTX-2.3 or other pipelines
  • Replacing unwanted original audio with clean Foley
  • Prototyping sound design quickly without sourcing stock libraries
  • Generating material-specific sounds where the visual contact is clear

It is not a substitute for dialogue synthesis, music generation, or final professional mixing. Use it for Foley and SFX layers.

Get Started

Recommended inference recipe:

  • Base model: LTX-2.3-22B
  • LoRA strength: 0.8–1.0
  • Guidance scale: 6.0
  • Inference steps: 30
  • Resolution: 960×544 @ 24 fps
  • Frame count: 1 + 8×n (89 frames is the main training bucket; tested up to 169)
  • STG: scale 1.0, block [29], mode stg_av
  • Checkpoint: step 500 (best loudness)

Prompting tip:

Write one to three dense sentences that name the action, material, acoustic character, and timing. End with: No speech is present. No music is present.

Example:

A sledgehammer smashes a glass bottle: a heavy whoosh, a sharp glass shatter and tinkling shards. No speech is present. No music is present.

Base negative prompt:

music, melody, song, singing, vocals, score, soundtrack, speech, dialogue, talking, narration, muffled, dull, distant, low-pass filtered, underwater, distorted, clipped, static, room tone only, silence, generic sound effect

Get the model and workflow:

Model: Lightricks/LTX-2.3-22b-LoRA-Foley-V2A on Hugging Face

ComfyUI nodes: ComfyUI-LTXVideo on GitHub

Base model: Lightricks/LTX-2.3 on Hugging Face

Credits

Demo footage sources:

  • Side-by-side comparison clips are from Pexels under the Pexels Free License.
  • Hero example clips were generated end-to-end with no stock footage: Seedream v5 Pro (text → still) → LTX-2.3 I2V (still → muted video) → Foley LoRA v4 (video → audio).
  • Publishing assets, ComfyUI workflows, and the model card are included in the Hugging Face release.

Acknowledgments:

  • Base model: Lightricks/LTX-2.3
  • Training infrastructure: LTX-2 Community Trainer
  • Dataset: FoleyBench (CC BY-NC-SA) and a curated Foley rebuild set

License: LTX-2 Community License. Dataset components may carry additional terms.

Table of Contents: