Fix API to match actual ltx-pipelines 1.1.5 interface

- Use DistilledPipeline constructor (not from_config)
- Use ImageConditioningInput namedtuple for image conditioning
- Use LoraPathStrengthAndSDOps for LoRA config
- Use QuantizationPolicy with fp8_cast sd_ops directly
- Update spatial_upsampler to required x2-1.1 model
- Update gemma_root to google/gemma-3-12b-it-qat-q4_0-unquantized
- Handle video output as Iterator[torch.Tensor] from pipeline
- Flesh out README with all model download commands, endpoint docs, env vars, frame/resolution tables
This commit is contained in:
2026-06-01 19:37:16 -04:00
parent 2d50679ec0
commit a33a3a593d
3 changed files with 261 additions and 100 deletions
+164 -67
View File
@@ -1,11 +1,19 @@
# Revids
REST API for generating video from images using [LTX-2.3](https://docs.ltx.video/open-source-model/integration-tools/pytorch-api).
REST API for generating video from images using [LTX-2.3](https://github.com/Lightricks/LTX-2) distilled pipeline.
## Architecture
- **FastAPI** REST server with auto-generated OpenAPI docs
- **LTX-2.3 DistilledPipeline** — two-stage video generation (8 steps stage 1, 4 steps stage 2)
- **Async task queue** — submit job, poll status, download when ready
- **SQLite** — job persistence via aiosqlite
- **Single GPU lock** — jobs serialized to avoid VRAM contention
## Quick Start
```bash
# 1. Clone the repo (including submodule)
# 1. Clone the repo (including LTX-2 submodule)
git clone --recursive <repo-url> revids
cd revids
@@ -13,138 +21,227 @@ cd revids
python -m venv .venv && source .venv/bin/activate
pip install -e libs/LTX-2/packages/ltx-core -e libs/LTX-2/packages/ltx-pipelines
# 3. Install API deps
# 3. Install API dependencies
pip install -r requirements.txt
# 4. Download LTX-2.3 model weights
# 4. Download model weights from HuggingFace
## Distilled model checkpoint
huggingface-cli download Lightricks/LTX-2.3 \
--include "ltx-2.3-22b-distilled-1.1.safetensors" \
--local-dir models/
# 5. Download Gemma/GPT-4o text encoder
huggingface-cli download Lightricks/LTX-2 \
--local-dir models/gpt-4o-5805-ava-gguf-model
# 6. Download spatial upsampler (optional, for higher-res output)
huggingface-cli download Lightricks/LTX-2 \
--include "ltx-2.3-22b-spatial-upscaler.safetensors" \
## Spatial upsampler (required for two-stage pipeline)
huggingface-cli download Lightricks/LTX-2.3 \
--include "ltx-2.3-spatial-upscaler-x2-1.1.safetensors" \
--local-dir models/
# 7. Run
## Gemma text encoder
huggingface-cli download google/gemma-3-12b-it-qat-q4_0-unquantized \
--local-dir models/gemma-3-12b-it-qat-q4_0-unquantized
# 5. Optimize memory (optional but recommended)
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
# 6. Run
uvicorn app.main:app --reload
```
The API will be available at `http://localhost:8000`. Interactive docs at `http://localhost:8000/docs`.
The API is available at `http://localhost:8000`. Swagger docs at `http://localhost:8000/docs`.
## API Endpoints
### Submit a Job
### Submit Job — `POST /generate`
Upload an image and generation parameters. Returns a job ID immediately.
```bash
curl -X POST http://localhost:8000/generate \
-F "image=@photo.jpg" \
-F "prompt=A cinematic pan across the landscape" \
-F "prompt=A cinematic pan right across a misty mountain range at dawn, golden light breaking through clouds" \
-F "width=768" \
-F "height=512" \
-F "num_frames=65" \
-F "fps=24.0"
-F "num_frames=97" \
-F "fps=24.0" \
-F "seed=42"
```
Response:
**Parameters:**
| Param | Type | Default | Description |
|---|---|---|---|
| `image` | file (required) | — | Input image (PNG, JPG, WebP, etc.) |
| `prompt` | string | `""` | Optional text prompt to guide generation |
| `width` | int | 768 | Output width (must be divisible by 32) |
| `height` | int | 512 | Output height (must be divisible by 32) |
| `num_frames` | int | 65 | Frame count (must be `8n+1`: 9, 17, …, 65, 97, 121, 161, 257) |
| `fps` | float | 24.0 | Frames per second |
| `seed` | int | random | Random seed for reproducibility |
**Response:**
```json
{"job_id": "abc123def456", "status": "pending"}
{
"job_id": "abc123def456",
"status": "pending",
"message": "Job submitted. Poll GET /jobs/{job_id} for status."
}
```
### Check Job Status
### Check Job Status — `GET /jobs/{job_id}`
```bash
curl http://localhost:8000/jobs/abc123def456
```
Response:
**Response:**
```json
{
"job_id": "abc123def456",
"status": "completed",
"prompt": "A cinematic pan across the landscape",
"params": {"width": 768, "height": 512, "num_frames": 65, "fps": 24.0},
"prompt": "A cinematic pan right...",
"params": {"width": 768, "height": 512, "num_frames": 97, "fps": 24.0, "seed": 42},
"error": null,
"created_at": "2025-01-01T00:00:00+00:00",
"completed_at": "2025-01-01T00:01:30+00:00"
"created_at": "2025-01-01T12:00:00+00:00",
"completed_at": "2025-01-01T12:01:30+00:00"
}
```
### Download Video (`.mp4`)
**Status values:** `pending``processing``completed` or `failed`
### Download Video — `GET /jobs/{job_id}/download`
Returns the `.mp4` video file when the job is completed. Returns 404 if not ready.
```bash
curl http://localhost:8000/jobs/abc123def456/download -o output.mp4
```
### Delete Job (removes DB record + video file)
### Delete Job — `DELETE /jobs/{job_id}`
Removes the job record and associated video file.
```bash
curl -X DELETE http://localhost:8000/jobs/abc123def456
```
### Health Check
### Health Check — `GET /health`
Returns GPU availability and device info.
```bash
curl http://localhost:8000/health
```
Returns GPU availability and device name.
## Generation Constraints
```json
{"status": "ok", "gpu": true, "gpu_name": "NVIDIA A100-SXM4-80GB"}
```
- **Width/Height**: must be divisible by 32 (min 256, max 2048)
- **num_frames**: must follow `8n+1` pattern (9, 17, ..., 65, 97, 121, 161, 257, ...)
- **fps**: 0 < fps <= 60
## Frame Count Reference
Num_frames must follow `8n+1` pattern. Common values:
| Frames | Duration (24 fps) | Duration (25 fps) |
|---|---|---|
| 65 | ~2.7s | ~2.6s |
| 97 | ~4.0s | ~3.9s |
| 121 | ~5.0s | ~4.8s |
| 161 | ~6.7s | ~6.4s |
| 257 | ~10.7s | ~10.3s |
## Resolution Reference
Width and height must be divisible by 32 (min 256, max 2048).
| Resolution | Aspect Ratio |
|---|---|
| 768×512 | 3:2 landscape |
| 512×768 | 2:3 portrait |
| 704×512 | 4:3 standard |
| 640×640 | 1:1 square |
## Configuration
All settings in `app/config.py` can be overridden via `.env` or env vars prefixed with `REVIDS_`:
All settings can be overridden via `.env` file (placed in project root) or environment variables prefixed with `REVIDS_`. See `app/config.py` for all available options.
| Variable | Default | Description |
| Env Variable | Default | Description |
|---|---|---|
| `REVIDS_HOST` | `0.0.0.0` | Bind address |
| `REVIDS_PORT` | `8000` | Port |
| `REVIDS_LTX_DISTILLED_CHECKPOINT` | `models/ltx-2.3-22b-distilled-1.1.safetensors` | Distilled model checkpoint path |
| `REVIDS_LTX_GEMMA_ROOT` | `models/gemma-3-12b-it-qat-q4_0-unquantized` | Gemma text encoder path |
| `REVIDS_LTX_SPATIAL_UPSAMPLER` | `models/ltx-2.3-spatial-upscaler-x2-1.1.safetensors` | Spatial upsampler checkpoint |
| `REVIDS_LTX_QUANTIZATION` | `fp8_cast` | Quantization: `fp8_cast`, `fp8_scaled_mm`, or `""` (bfloat16) |
| `REVIDS_LTX_LORAS` | *(empty)* | LoRA config: `"path1:1.0,path2:0.5"` |
| `REVIDS_MAX_CONCURRENT_JOBS` | `1` | Concurrent GPU jobs |
| `REVIDS_PORT` | `8000` | Server port |
| `REVIDS_RELOAD` | `true` | uvicorn auto-reload |
| `REVIDS_LTX_DISTILLED_CHECKPOINT` | `models/ltx-2.3-22b-distilled-1.1.safetensors` | Distilled model path |
| `REVIDS_LTX_GEMMA_ROOT` | `models/gpt-4o-5805-ava-gguf-model` | Gemma text encoder path |
| `REVIDS_LTX_SPATIAL_UPSAMPLER` | *(none)* | Spatial upsampler model path |
| `REVIDS_LTX_QUANTIZATION` | `fp8_cast` | Quantization mode (`fp8_cast`, `fp8_scaled_mm`, or unset for bfloat16) |
| `REVIDS_LTX_LORAS` | *(empty)* | LoRA paths, colon-separated: `"models/lora1.safetensors:1.0,models/lora2.safetensors:0.5"` |
| `REVIDS_MAX_CONCURRENT_JOBS` | `1` | GPU concurrency |
### LoRA Support
To load custom or Lightricks LoRAs:
```bash
# Single LoRA at full strength
export REVIDS_LTX_LORAS="models/my-style.safetensors:1.0"
# Multiple LoRAs
export REVIDS_LTX_LORAS="models/style.safetensors:1.0,models/motion.safetensors:0.5"
```
Lightricks-published LoRAs:
- [IC-LoRA Union Control](https://huggingface.co/Lightricks/LTX-2.3-22b-IC-LoRA-Union-Control)
- [IC-LoRA Motion Track Control](https://huggingface.co/Lightricks/LTX-2.3-22b-IC-LoRA-Motion-Track-Control)
- [IC-LoRA LipDub](https://huggingface.co/Lightricks/LTX-2.3-22b-IC-LoRA-LipDub)
- Camera Control LoRAs (dolly in/out, jib up/down, static, etc.)
## Memory Optimization
```bash
# Improve CUDA memory allocation fragmentation
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
# Use FP8 quantization (~40% VRAM reduction, minimal quality loss)
export REVIDS_LTX_QUANTIZATION=fp8_cast
# For Hopper GPUs (H100) with TensorRT-LLM:
# export REVIDS_LTX_QUANTIZATION=fp8_scaled_mm
```
## Hardware Requirements
- **GPU**: NVIDIA with CUDA support
- **VRAM**: ≥ 24 GB (bfloat16), ~14 GB with FP8 quantization
- **Disk**: ~30 GB for all model weights (distilled + upsampler + Gemma encoder)
## Project Structure
```
revids/
app/
main.py # FastAPI app + routes
config.py # Settings (pydantic-settings)
models.py # Request/response schemas
service.py # LTX pipeline wrapper + job processing
database.py # SQLite job store
__init__.py
main.py # FastAPI app + HTTP routes
config.py # Settings (pydantic-settings, env var support)
models.py # Request/response Pydantic models
service.py # LTX pipeline wrapper, job processing, video encoding
database.py # SQLite job store (aiosqlite)
libs/
LTX-2/ # LTX-2 git submodule
LTX-2/ # LTX-2 git submodule (source for ltx-core/ltx-pipelines)
videos/ # Generated .mp4 output (gitignored)
requirements.txt
videos/ # Generated video output (gitignored)
.gitignore
```
## Hardware Requirements
## Prompting Guide
- NVIDIA GPU with CUDA (>= 24GB VRAM for bfloat16, ~14GB with FP8)
- ~30GB disk for model weights
For best results with LTX-2.3, write detailed, chronological motion descriptions:
## Memory Optimization
1. Start with the main action in one sentence
2. Add specific movement details and gestures
3. Describe appearances precisely
4. Include background and environment
5. Specify camera angle and movement
6. Describe lighting and color palette
7. Keep under 200 words
```bash
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
```
Full guide: [https://ltx.video/blog/how-to-prompt-for-ltx-2](https://ltx.video/blog/how-to-prompt-for-ltx-2)
Set `REVIDS_LTX_QUANTIZATION=fp8_cast` (~40% VRAM reduction) or `fp8_scaled_mm` (Hopper GPUs only).
## License
## LoRA Support
Configure LoRAs via environment variable:
```bash
export REVIDS_LTX_LORAS="models/my-style.safetensors:1.0,models/my-motion.safetensors:0.5"
```
Or set as a comma-separated list of `path:scale` entries.
See [LTX-2 LICENSE](libs/LTX-2/LICENSE) for model usage terms.