Building foundation models for the physical world.

Our research team builds open world models that can generate, simulate, and shape the world.

LTX-2.3 Examples Video

LTX-2 : Efficient Joint Audio-Visual Foundation Model

Asymmetric dual-stream transformer (14B video + 5B audio) generating synchronized audio-visual content in a unified pass, open-sourced.

Foundational research

Publications

Build on LTX

Independent research powered by our open foundations. 

Robotics

Imagination, navigation, and world-model research powering perception and planning for embodied agents operating in the real world.

A modular navigation paradigm that imagines future video trajectories from language subgoals, then extracts actions via inverse dynamics — no robot demos required.

arxiv

A surgical action world model conditioned on lightweight signals — language, reference scene, affordance, tool-tip trajectories — for controllable laparoscopic video generation.

Turns video diffusion models into counterfactual world models by conditioning on digital-twin scene representations modified by an LLM.

A unified video-generative platform for robotic manipulation that combines policy learning, action-conditioned simulation, and standardized evaluation.

Medical

Surgical and clinical world models for operating-room event recognition and procedural understanding, supporting training, monitoring, and analysis.

An OR video diffusion framework that synthesizes routine and rare operating-room events from abstract geometric representations to support ambient surgical intelligence.

Diffusion-based world models fine-tuned on laparoscopic footage to simulate the biomechanics of robotic surgical suturing with high temporal fidelity.

Creative Applications

Editing, compositing, and controllable generation that extends LTX into expressive media workflows, giving creators finer control over what's generated and how it's refined.

OmniForcing

March 12, 2026
arxiv

Distills an offline dual-stream audio-visual diffusion model into a real-time streaming generator running at ~25 FPS with synchronized multimodal output.

OmniEdit

March 10, 2026
arxiv

A training-free framework for lip synchronization and audio-visual editing that stabilizes the editing trajectory by removing stochastic elements.

mAVE

March 7, 2026
arxiv

A watermarking framework that cryptographically binds audio and video latents in joint generation models, blocking deepfake swap attacks with >99% integrity.

FREE-Edit

March 1, 2026
arxiv

A zero-shot, training-free image-driven video editing method for rectified flow models that modulates injection intensity using an editing-aware mask.

PropFly

February 24, 2026
arxiv

A propagation-based video editing pipeline that learns from on-the-fly supervision generated by pre-trained video diffusion models — no paired data needed.

EditCtrl

February 16, 2026
arxiv

An efficient video inpainting control framework that focuses compute on masked tokens, making localized edits 10x cheaper without sacrificing quality.

HairWeaver

February 11, 2026
arxiv

A few-shot, sim-to-real video diffusion pipeline that animates realistic, expressive hair motion from a single human image using lightweight LoRA modules.

What Happens Next?

December 15, 2025
arxiv

A unified Qwen-VL + LTX framework for next-scene prediction, trained with a causal consistency reward to anticipate plausible futures.

DisMo

November 28, 2025
arxiv

Learns appearance-independent motion representations from raw video, enabling open-world motion transfer between semantically unrelated entities.

READ

August 5, 2025
arxiv

A real-time DiT talking-head framework using a temporal VAE and Speech Autoencoder for efficient, well-aligned audio-driven synthesis.

EraserDiT

June 15, 2025
arxiv

A fast DiT-based video inpainting method using a Circular Position-Shift strategy for strong long-term temporal consistency on large masks.

OmnimatteZero

March 23, 2025
arxiv

A training-free omnimatte approach using pre-trained video diffusion models to decompose videos into object layers and effects at real-time speed.

VACE

March 10, 2025
arxiv

An all-in-one video creation and editing framework that unifies reference-to-video generation, video-to-video editing, and masked editing through a single Video Condition Unit interface.

Gaming & 3D

World-building and 3D scene generation for interactive, game-ready environments — from open worlds to indoor scenes, conditioned on layout and gameplay intent.

Activates the in-context generation ability of large video models to produce multiple viewpoint-consistent videos of the same shared world.

The first open platform that benchmarks generative world models inside closed-loop environments, measuring whether they actually help embodied agents succeed.

A world model that pairs panoramic video generation with an evolving explicit 3D memory for spatially consistent long-horizon exploration.

Autonomous Driving

Driving-world simulation and data augmentation for safer perception, synthesizing diverse scenarios to stress-test autonomous stacks.

arxiv

Decouples motion from appearance to adapt generalist video diffusion models into controllable driving world models with under 6% of prior compute.

Unifies video generation and motion planning by feeding the video world model's latents directly into a diffusion planner for consistent driving trajectories.

Defense

Generative modeling for synthetic data and adversarial robustness, providing high-fidelity environments where real-world data is limited or sensitive.

Uses video diffusion models to synthesize realistic 3D camera and scene-motion variations, augmenting scarce datasets such as UAV imagery.

Research in Conversation

Research shared through talks, podcasts, workshops, and public conversations from the LTX Research team.

session

Orchestrating Multimodal Control for Unified Audio-Visual Synthesis

Naomi Ken Korem
Computer Vision Researcher, Lightricks
An upcoming SIGGRAPH 2026 session on orchestrating multimodal controls to unify audio-visual synthesis, presented by Naomi Ken Korem and Urska Jelercic.
View Session
Source:
siggraph
Source:
siggraph
Jul 21st, 2026
video

I'm Putting Real-Time AI Video on Your Phone

Zeev Farbman
Co-founder & CEO
Zeev Farbman joins the Almost Human podcast to talk about bringing real-time AI video generation onto mobile devices.
Watch Now
Source:
Almost Human
Source:
Almost Human
Nov 25th, 2025
podcast

How to Structure a Growing AI Company

Yaron Inger
Co-founder & CTO
Yaron Inger on the practical challenges of structuring and scaling a fast-growing AI company, from co-founding Lightricks to leading its engineering.
Listen Now
Source:
IT Labs
Source:
IT Labs
Jan 24th, 2025
video

Deep Dive into AI for Video and Its Consequences

Zeev Farbman
Co-founder & CEO
Zeev Farbman's deep-dive conversation on the state of AI video and its broader consequences for filmmaking and creators.
Watch Now
Source:
CineD
Source:
CineD
Aug 21st, 2024
article

Interview Series: Zeev Farbman

Zeev Farbman
Co-founder & CEO, Lightricks
An Unite.AI interview with Zeev Farbman on co-founding and leading Lightricks, and the research vision behind LTX.
Read Now
Source:
Unite.AI
Source:
Unite.AI
Jan 24th, 2024
//

Academic Programs

Build with us

Whether you're just getting started or building the next generation of physical AI, we're here to support you.