arXiv cs.CL
7/16/2026

Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation
Short summary
Text2Sign is a text-conditioned diffusion model for generating short sign-language video clips that runs on a single NVIDIA L4 GPU. It combines a frozen vision-language text encoder with a 3D encoder-decoder and factorized spatiotemporal attention to reduce computational cost. The system generates 32-frame 64x64 clips at 2.54 FPS with 3.12 GB peak memory, but shows only weak prompt sensitivity and lacks expert linguistic evaluation. The authors position it as a single-GPU research baseline rather than a production system, with code publicly available.
- •Text2Sign generates short sign-language clips on a single L4 GPU with 3.12 GB peak memory
- •Frozen text conditioning improves validation loss but prompt-specific separation remains limited
- •Positioned as a research baseline; restricted to low-resolution, short clips without linguistic evaluation
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


