Our research team builds open world models that can generate, simulate, and shape the world.

Conditions a text-to-audio flow-matching model on multiple reference voices and a free-form prompt to generate natural multi-speaker audio scenes with real-world ambient texture.
A single-model approach that adapts a joint audio-visual diffusion foundation model for video dubbing via a lightweight LoRA, replacing complex task-specific pipelines.
Shows HDR video generation can be done simply by applying a logarithmic encoding to pretrained generative priors, avoiding new representations and extra training data.
A lightweight, extendable framework built on LTX-2 that trains each video and audio control modality — depth, pose, camera, audio — as a separate LoRA, with no architectural changes.
A practical SIGGRAPH Asia course guiding developers through fine-tuning state-of-the-art text-to-video models with LoRA.
A transformer-based latent diffusion model that tightly integrates the Video-VAE and denoiser, reaching a 1:192 compression ratio to generate video faster than real time — with open weights.
Independent research powered by our open foundations.

Imagination, navigation, and world-model research powering perception and planning for embodied agents operating in the real world.
A modular navigation paradigm that imagines future video trajectories from language subgoals, then extracts actions via inverse dynamics — no robot demos required.
A surgical action world model conditioned on lightweight signals — language, reference scene, affordance, tool-tip trajectories — for controllable laparoscopic video generation.
Turns video diffusion models into counterfactual world models by conditioning on digital-twin scene representations modified by an LLM.
A unified video-generative platform for robotic manipulation that combines policy learning, action-conditioned simulation, and standardized evaluation.

Surgical and clinical world models for operating-room event recognition and procedural understanding, supporting training, monitoring, and analysis.
An OR video diffusion framework that synthesizes routine and rare operating-room events from abstract geometric representations to support ambient surgical intelligence.
Diffusion-based world models fine-tuned on laparoscopic footage to simulate the biomechanics of robotic surgical suturing with high temporal fidelity.

Editing, compositing, and controllable generation that extends LTX into expressive media workflows, giving creators finer control over what's generated and how it's refined.
Distills an offline dual-stream audio-visual diffusion model into a real-time streaming generator running at ~25 FPS with synchronized multimodal output.
A training-free framework for lip synchronization and audio-visual editing that stabilizes the editing trajectory by removing stochastic elements.
A watermarking framework that cryptographically binds audio and video latents in joint generation models, blocking deepfake swap attacks with >99% integrity.
A zero-shot, training-free image-driven video editing method for rectified flow models that modulates injection intensity using an editing-aware mask.
A propagation-based video editing pipeline that learns from on-the-fly supervision generated by pre-trained video diffusion models — no paired data needed.
An efficient video inpainting control framework that focuses compute on masked tokens, making localized edits 10x cheaper without sacrificing quality.
A few-shot, sim-to-real video diffusion pipeline that animates realistic, expressive hair motion from a single human image using lightweight LoRA modules.
A unified Qwen-VL + LTX framework for next-scene prediction, trained with a causal consistency reward to anticipate plausible futures.
Learns appearance-independent motion representations from raw video, enabling open-world motion transfer between semantically unrelated entities.
A real-time DiT talking-head framework using a temporal VAE and Speech Autoencoder for efficient, well-aligned audio-driven synthesis.
A fast DiT-based video inpainting method using a Circular Position-Shift strategy for strong long-term temporal consistency on large masks.
A training-free omnimatte approach using pre-trained video diffusion models to decompose videos into object layers and effects at real-time speed.
An all-in-one video creation and editing framework that unifies reference-to-video generation, video-to-video editing, and masked editing through a single Video Condition Unit interface.

World-building and 3D scene generation for interactive, game-ready environments — from open worlds to indoor scenes, conditioned on layout and gameplay intent.
Activates the in-context generation ability of large video models to produce multiple viewpoint-consistent videos of the same shared world.
The first open platform that benchmarks generative world models inside closed-loop environments, measuring whether they actually help embodied agents succeed.
A world model that pairs panoramic video generation with an evolving explicit 3D memory for spatially consistent long-horizon exploration.

Driving-world simulation and data augmentation for safer perception, synthesizing diverse scenarios to stress-test autonomous stacks.
Decouples motion from appearance to adapt generalist video diffusion models into controllable driving world models with under 6% of prior compute.
Unifies video generation and motion planning by feeding the video world model's latents directly into a diffusion planner for consistent driving trajectories.

Generative modeling for synthetic data and adversarial robustness, providing high-fidelity environments where real-world data is limited or sensitive.
Uses video diffusion models to synthesize realistic 3D camera and scene-motion variations, augmenting scarce datasets such as UAV imagery.
Research shared through talks, podcasts, workshops, and public conversations from the LTX Research team.
Academic Programs
Whether you're just getting started or building the next generation of physical AI, we're here to support you.
Giving builders early access, technical resources, and collaboration opportunities.
Helping developers build for robotics, automation, manufacturing, and beyond.
Supporting universities, research labs, and students advancing world models.