- Use DistilledPipeline constructor (not from_config) - Use ImageConditioningInput namedtuple for image conditioning - Use LoraPathStrengthAndSDOps for LoRA config - Use QuantizationPolicy with fp8_cast sd_ops directly - Update spatial_upsampler to required x2-1.1 model - Update gemma_root to google/gemma-3-12b-it-qat-q4_0-unquantized - Handle video output as Iterator[torch.Tensor] from pipeline - Flesh out README with all model download commands, endpoint docs, env vars, frame/resolution tables
Revids
REST API for generating video from images using LTX-2.3 distilled pipeline.
Architecture
- FastAPI REST server with auto-generated OpenAPI docs
- LTX-2.3 DistilledPipeline — two-stage video generation (8 steps stage 1, 4 steps stage 2)
- Async task queue — submit job, poll status, download when ready
- SQLite — job persistence via aiosqlite
- Single GPU lock — jobs serialized to avoid VRAM contention
Quick Start
# 1. Clone the repo (including LTX-2 submodule)
git clone --recursive <repo-url> revids
cd revids
# 2. Create venv and install LTX packages (editable, from submodule)
python -m venv .venv && source .venv/bin/activate
pip install -e libs/LTX-2/packages/ltx-core -e libs/LTX-2/packages/ltx-pipelines
# 3. Install API dependencies
pip install -r requirements.txt
# 4. Download model weights from HuggingFace
## Distilled model checkpoint
huggingface-cli download Lightricks/LTX-2.3 \
--include "ltx-2.3-22b-distilled-1.1.safetensors" \
--local-dir models/
## Spatial upsampler (required for two-stage pipeline)
huggingface-cli download Lightricks/LTX-2.3 \
--include "ltx-2.3-spatial-upscaler-x2-1.1.safetensors" \
--local-dir models/
## Gemma text encoder
huggingface-cli download google/gemma-3-12b-it-qat-q4_0-unquantized \
--local-dir models/gemma-3-12b-it-qat-q4_0-unquantized
# 5. Optimize memory (optional but recommended)
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
# 6. Run
uvicorn app.main:app --reload
The API is available at http://localhost:8000. Swagger docs at http://localhost:8000/docs.
API Endpoints
Submit Job — POST /generate
Upload an image and generation parameters. Returns a job ID immediately.
curl -X POST http://localhost:8000/generate \
-F "image=@photo.jpg" \
-F "prompt=A cinematic pan right across a misty mountain range at dawn, golden light breaking through clouds" \
-F "width=768" \
-F "height=512" \
-F "num_frames=97" \
-F "fps=24.0" \
-F "seed=42"
Parameters:
| Param | Type | Default | Description |
|---|---|---|---|
image |
file (required) | — | Input image (PNG, JPG, WebP, etc.) |
prompt |
string | "" |
Optional text prompt to guide generation |
width |
int | 768 | Output width (must be divisible by 32) |
height |
int | 512 | Output height (must be divisible by 32) |
num_frames |
int | 65 | Frame count (must be 8n+1: 9, 17, …, 65, 97, 121, 161, 257) |
fps |
float | 24.0 | Frames per second |
seed |
int | random | Random seed for reproducibility |
Response:
{
"job_id": "abc123def456",
"status": "pending",
"message": "Job submitted. Poll GET /jobs/{job_id} for status."
}
Check Job Status — GET /jobs/{job_id}
curl http://localhost:8000/jobs/abc123def456
Response:
{
"job_id": "abc123def456",
"status": "completed",
"prompt": "A cinematic pan right...",
"params": {"width": 768, "height": 512, "num_frames": 97, "fps": 24.0, "seed": 42},
"error": null,
"created_at": "2025-01-01T12:00:00+00:00",
"completed_at": "2025-01-01T12:01:30+00:00"
}
Status values: pending → processing → completed or failed
Download Video — GET /jobs/{job_id}/download
Returns the .mp4 video file when the job is completed. Returns 404 if not ready.
curl http://localhost:8000/jobs/abc123def456/download -o output.mp4
Delete Job — DELETE /jobs/{job_id}
Removes the job record and associated video file.
curl -X DELETE http://localhost:8000/jobs/abc123def456
Health Check — GET /health
Returns GPU availability and device info.
curl http://localhost:8000/health
{"status": "ok", "gpu": true, "gpu_name": "NVIDIA A100-SXM4-80GB"}
Frame Count Reference
Num_frames must follow 8n+1 pattern. Common values:
| Frames | Duration (24 fps) | Duration (25 fps) |
|---|---|---|
| 65 | ~2.7s | ~2.6s |
| 97 | ~4.0s | ~3.9s |
| 121 | ~5.0s | ~4.8s |
| 161 | ~6.7s | ~6.4s |
| 257 | ~10.7s | ~10.3s |
Resolution Reference
Width and height must be divisible by 32 (min 256, max 2048).
| Resolution | Aspect Ratio |
|---|---|
| 768×512 | 3:2 landscape |
| 512×768 | 2:3 portrait |
| 704×512 | 4:3 standard |
| 640×640 | 1:1 square |
Configuration
All settings can be overridden via .env file (placed in project root) or environment variables prefixed with REVIDS_. See app/config.py for all available options.
| Env Variable | Default | Description |
|---|---|---|
REVIDS_LTX_DISTILLED_CHECKPOINT |
models/ltx-2.3-22b-distilled-1.1.safetensors |
Distilled model checkpoint path |
REVIDS_LTX_GEMMA_ROOT |
models/gemma-3-12b-it-qat-q4_0-unquantized |
Gemma text encoder path |
REVIDS_LTX_SPATIAL_UPSAMPLER |
models/ltx-2.3-spatial-upscaler-x2-1.1.safetensors |
Spatial upsampler checkpoint |
REVIDS_LTX_QUANTIZATION |
fp8_cast |
Quantization: fp8_cast, fp8_scaled_mm, or "" (bfloat16) |
REVIDS_LTX_LORAS |
(empty) | LoRA config: "path1:1.0,path2:0.5" |
REVIDS_MAX_CONCURRENT_JOBS |
1 |
Concurrent GPU jobs |
REVIDS_PORT |
8000 |
Server port |
REVIDS_RELOAD |
true |
uvicorn auto-reload |
LoRA Support
To load custom or Lightricks LoRAs:
# Single LoRA at full strength
export REVIDS_LTX_LORAS="models/my-style.safetensors:1.0"
# Multiple LoRAs
export REVIDS_LTX_LORAS="models/style.safetensors:1.0,models/motion.safetensors:0.5"
Lightricks-published LoRAs:
- IC-LoRA Union Control
- IC-LoRA Motion Track Control
- IC-LoRA LipDub
- Camera Control LoRAs (dolly in/out, jib up/down, static, etc.)
Memory Optimization
# Improve CUDA memory allocation fragmentation
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
# Use FP8 quantization (~40% VRAM reduction, minimal quality loss)
export REVIDS_LTX_QUANTIZATION=fp8_cast
# For Hopper GPUs (H100) with TensorRT-LLM:
# export REVIDS_LTX_QUANTIZATION=fp8_scaled_mm
Hardware Requirements
- GPU: NVIDIA with CUDA support
- VRAM: ≥ 24 GB (bfloat16), ~14 GB with FP8 quantization
- Disk: ~30 GB for all model weights (distilled + upsampler + Gemma encoder)
Project Structure
revids/
app/
__init__.py
main.py # FastAPI app + HTTP routes
config.py # Settings (pydantic-settings, env var support)
models.py # Request/response Pydantic models
service.py # LTX pipeline wrapper, job processing, video encoding
database.py # SQLite job store (aiosqlite)
libs/
LTX-2/ # LTX-2 git submodule (source for ltx-core/ltx-pipelines)
videos/ # Generated .mp4 output (gitignored)
requirements.txt
.gitignore
Prompting Guide
For best results with LTX-2.3, write detailed, chronological motion descriptions:
- Start with the main action in one sentence
- Add specific movement details and gestures
- Describe appearances precisely
- Include background and environment
- Specify camera angle and movement
- Describe lighting and color palette
- Keep under 200 words
Full guide: https://ltx.video/blog/how-to-prompt-for-ltx-2
License
See LTX-2 LICENSE for model usage terms.